Key concepts
The three components
ContextBase
Flume’s product for understanding and answering from your data estate. ContextBase connects to your systems, extracts metadata and artifacts, and synthesizes them into a knowledge graph: a navigable map of what your data means. It does this through three stages:
- Context Generation
- Extracting and synthesizing raw metadata into meaningful relationships.
- Context Management
- Resolving conflicts between sources, scoring confidence, and versioning changes.
- Context Storage
- Persisting everything in a graph database where assets are nodes and relationships are edges.
When you ask “what does this column mean?” in Chat, or with /explore in the CLI, ContextBase guides the LLM with an iterative loop of vector and Cypher search and validation. For the claim line’s allowed amount, that meaning comes from the schema definition, the stored procedure that sets the value, the wiki page that documents the lesser-of rule, and the dashboard that reads the column.
Lakehouse (Trino + Iceberg)
The analytical engine. When either product needs to profile data distributions, validate pipeline outputs, or run cross-system analysis, it uses the Lakehouse. Trino is the query engine: distributed SQL that can query data in place without moving it. Iceberg is the table format: ACID transactions, schema evolution, and time travel on top of object storage.
You don’t interact with the Lakehouse directly. It powers the profiling, validation, and analytical capabilities that show up in your workflows.
Data Pipelines
The integration platform. Source endpoints pull data from origin systems. Destination endpoints push to targets. A canonical data lake in the middle normalizes everything. Each side has its own protocol, schedule, and mapping and transform logic, configured independently.
One source can feed many destinations. Many sources can feed one destination. Any protocol to any protocol. The canonical layer in the middle means adding a new destination doesn’t require touching the source.
How they relate
Connectors are shared: configure a connection once, and both products use it. ContextBase tells Data Pipelines where to pull data from, how to join it, and how to validate it. Pipeline execution generates metadata that feeds back into ContextBase.
What you work with
The Application organizes ContextBase by the work you do.
- Data
- Browse every server, table, and row. Profile any column.
- Chat
- Ask in plain language. Every answer arrives with its receipts and its answer checks.
- Value Lab
- The full library of payer value cases, checked against your estate automatically. Each case reads Ready, Gaps, or Blocked, and names the data it needs.
- Dashboards
- Built once with the model. A refresh re-runs the saved queries without the model, so it spends no tokens.
- Apps
- Real tools and algorithms, such as plan benefit clustering. Each app declares what it reads, its inputs, and what it acts on.
- Bots
- Bots run on a schedule, read only by default, and deliver each run’s results inside ContextBase.
Glossary
- Data estate
- All systems where your organization’s data lives, moves, and gets transformed.
- ContextBase
- Flume’s product for understanding and answering from your data estate. It holds the knowledge graph and runs its analysis on the Lakehouse.
- Knowledge graph
- A synthesized, scored, and versioned map of what your data means, built from scans of your systems.
- Context scoring
- A confidence measure for each piece of context. Schema introspection scores higher than a stale Confluence page.
- Data asset
- Any discrete data object: table, view, column, stored procedure, dashboard, pipeline, or API endpoint.
- Lineage
- The chain of transformations connecting data assets. Column-level lineage traces individual fields; table-level lineage traces datasets.
- Upstream and downstream
- Relative position in lineage. A source table is upstream of the view that reads it.
- Canonical data lake
- The intermediate normalization layer in Data Pipelines, where data is converted to a common format.
- PAM session
- Privileged Access Management session. An approved, time-limited window of access to your data systems.
- Chat session
- A conversation with ContextBase in the Application or the CLI. Independent of PAM sessions.
- Answer check
- A named test run on an answer. It comes back met, worth a look, or unmet, and says so when it could not check.
- Receipt
- The record that comes with an answer: queries run, rows returned, sources read, and what was checked and what could not be.
- Approved finding
- A fact about your estate that a person approved. It lives in the knowledge layer, attached to its table.
- Knowledge layer
- The approved findings, each attached to the table it describes. Every answer that reads a table carries its findings, and the checks read them.
- Tiered analysis
- ContextBase’s framework for analysis depth: Tier 1 (deep introspection), Tier 2 (structured metadata), and Tier 3 (documentation linking).
- Connector
- A configured connection to a specific system in the data estate.
- Workflow
- A structured capability (Explore, Investigate, Create) accessed via slash commands.