Connection guide
What Flume needs to connect to each system in your data estate, and how we validate that every connection works.
How to use this guide
The requirements for every connector follow the same four-section structure:
- Required information
- The exact fields we need. No optional boilerplate: every field exists because we need it.
- Network considerations
- How to get packets from Flume to the target system, across every deployment model we’ve seen.
- Credential and auth management
- Auth method preference hierarchy (federation → secrets manager → secure transfer), session handling, and rotation.
- Validation checks
- The specific commands we run to confirm end-to-end connectivity. Four stages: reachability → authentication → permissions → functional test.
Connector categories
| Category | Systems | Why we connect |
|---|---|---|
| Databases, lakes, and warehouses | SQL Server, PostgreSQL, Snowflake, BigQuery, Redshift, Databricks, MongoDB, Oracle, MySQL, Teradata, SAP HANA, Cassandra, DynamoDB, Elasticsearch, CockroachDB, Vertica, Delta and Iceberg on object storage | Where structured data lives |
| Document and file repositories | SharePoint, Google Drive, Confluence, Notion, Box, S3, GCS, Azure Blob, SFTP, and network shares | Where unstructured knowledge lives, which may inform the context of structured data |
| Code repositories | GitHub, GitLab, Azure DevOps, Bitbucket | Where data logic lives |
| Decision and work tracking | Jira, Azure Boards, ServiceNow, Linear, Asana | Where data decisions get documented |
| Data catalogs and metadata | Alation, Collibra, Microsoft Purview, Atlan, DataHub, OpenMetadata | Where data documentation lives |
| BI and analytics | Tableau, Power BI, Looker, Metabase, Mode, Sigma | Where data gets consumed (reverse lineage) |
| ETL/ELT and orchestration | dbt, Airflow, Informatica, SSIS, Azure Data Factory, AWS Glue, Fivetran, and Airbyte metadata | Where transformation logic lives |
| API and application platforms | Salesforce, SAP | Where business logic and data intersect |
Pre-engagement checklist
Before configuring any connector, we need the following from your infrastructure and security teams:
- Network topology
- Are target systems reachable via public IP, VPN, VPC peering, PrivateLink, Private Endpoint, or SSH bastion? Include a network diagram if available.
- Firewall and IP allowlist
- Flume’s egress IPs, provided per environment, must be allowlisted.
- Credential delivery method
- Preferred approach: secrets manager integration (Vault, AWS Secrets Manager, Azure Key Vault), a 1Password shared vault, or Doppler.
- Auth policy
- Does your organization require OAuth/OIDC, SAML SSO for service accounts, or key-pair auth, or is password auth acceptable? Are there MFA requirements on service accounts?
- Session and timeout policy
- Maximum session duration, idle timeout, and concurrent session limits.
- Change control process
- Do firewall changes, new service accounts, or VPN configs require CAB approval? What are typical lead times?
- SSL/TLS requirements
- Is TLS mandatory? Self-signed or internal CA certificates? Minimum TLS version?
- Environment inventory
- All environments (dev, staging, prod) and whether each needs connectivity. Separate credentials per environment?
Auth preference hierarchy
For every connector, we prefer authentication methods in this order:
01Identity federation
OAuth service principals, IAM role assumption, key-pair auth. No secret crosses the wire.
02Secrets manager integration
You store credentials in your Vault, AWS Secrets Manager, or Azure Key Vault, and grant Flume’s service identity read access. We pull them at runtime.
03Secure credential transfer
1Password shared vault, Doppler, or Infisical. Credentials land in Flume’s secrets manager immediately; the transfer mechanism is ephemeral.
The goal is to never handle raw credentials when avoidable.
Universal validation protocol
Every connection goes through four stages before production sign-off:
01Network reachability
TCP/HTTPS connectivity to the target host:port from Flume infrastructure.
On failure: Firewall rule, VPN config, or DNS resolution issue → your network team.
02Authentication
Credential acceptance and identity confirmation (
SELECT 1or equivalent).On failure: Wrong credentials, expired password, or auth mechanism mismatch → your DBA or IAM team.
03Permissions
Schema, table, view, and procedure visibility matches the agreed scope.
On failure: Missing GRANTs → we provide the exact GRANT commands.
04Functional test
Representative read query, stored procedure execution, and result set structure validation.
On failure: Row-level security, column masking, or VPD issues → we adjust policies with you.
Each step is logged and timestamped, and you receive a validation report.
Enterprise networking patterns
| Pattern | What we need | Typical setup time |
|---|---|---|
| Site-to-site VPN | VPN gateway endpoint, IKE/IPSec parameters, PSK or certificate, Flume’s public IP | 1 to 2 weeks |
| SSH bastion | Bastion IP, SSH port, service account, public key exchange, target host list | 1 to 3 days |
| AWS PrivateLink | VPC endpoint service name, DNS, Flume’s AWS account ID | 3 to 5 days |
| Azure Private Endpoint | Private endpoint resource ID, DNS zone config, Flume’s Azure subscription ID | 3 to 5 days |
| GCP VPC peering | GCP project ID, VPC network name, non-overlapping IP ranges | 3 to 5 days |
| IP allowlisting | Flume provides static egress IPs; you add them to your firewall, security group, or ACL | 1 to 2 days |
| Zero Trust / BeyondCorp | IdP details for Flume’s service account, proxy or gateway URL, device trust requirements | 1 to 2 weeks |
Tools you already run
If you already run any of these, tell us during setup. Each one shortens the work or keeps your existing controls in place.
| If you use | How it helps | What to share |
|---|---|---|
| dbt | manifest.json gives column-level lineage without extra work. | The manifest.json from dbt Cloud or from your dbt Core repository. |
| HashiCorp Vault | Credentials stay in your Vault, and Flume pulls them at runtime. | Read access for Flume’s service identity. |
| Teleport or Boundary | Flume connects through your access proxy, not around it. | The proxy or gateway URL, and access for Flume’s service account. |
| Airflow or Astronomer | Scheduled jobs can be orchestrated in your Airflow once connections are live. | Nothing extra. Mention it during setup. |
Replication tools such as Fivetran, Airbyte, and Stitch copy data into a warehouse. They don’t provide live query access, stored procedure execution, or schema introspection, so Flume does not connect through them. If you use one, we connect to the warehouse it loads.