Google BigQuery
GCP native.
Access mode: Read-only or read-write
Trino connector. Queryable via Flume’s Lakehouse. This system can be accessed both through its native protocol (for metadata introspection) and via Trino federation (for data profiling and cross-system analytical queries).
Required information
| Field | Details |
|---|---|
| GCP project ID | Project(s) containing the datasets. |
| Dataset(s) | All datasets requiring access. |
| Service account | Dedicated GCP service account (e.g., flume-connector@<project>.iam.gserviceaccount.com). Provide JSON key file or configure Workload Identity Federation. |
| IAM roles | Read-only: BigQuery Data Viewer + BigQuery Job User. Read-write: BigQuery Data Editor + BigQuery Job User. |
| Billing project | If different from data project. Flume needs bigquery.jobs.create on billing project. |
| Location/region | Dataset region (US, EU, etc.). Cross-region queries have cost/latency implications. |
Network considerations
Default: Fully managed SaaS. Connections over HTTPS to googleapis.com. No VPN/peering needed.
VPC Service Controls: If org uses VPC-SC, add Flume’s service account project to the access perimeter or create an ingress rule.
No IP allowlisting needed unless VPC-SC is enforced.
Credential and auth management
Preferred: Workload Identity Federation. Flume federates through AWS or Azure identity. No long-lived keys. Requires Workload Identity Pool configuration on client side.
Acceptable: Service account JSON key. Standard approach. Flume stores it encrypted at rest in secrets manager.
OAuth scopes: https://www.googleapis.com/auth/bigquery (scoped to read-only or read-write).
No session concept: Each API call is independently authenticated. Token refresh automatic.
Stored procedure and logic access
BigQuery routines (procedures via CALL, functions via SELECT) require BigQuery Data Viewer to list and appropriate execute permissions. JavaScript UDFs need no special config. Remote functions backed by Cloud Functions require that the function is also accessible to Flume’s service account.
Validation checks
| Check | Method | Expected result |
|---|---|---|
| Authentication | bq query --project_id=<project> 'SELECT 1' | Returns 1 |
| Dataset access | bq ls <project>:<dataset> | Lists tables |
| Table read | SELECT * FROM <project>.<dataset>.<table> LIMIT 1 | Returns a row |
| Routine listing | SELECT routine_name FROM <dataset>.INFORMATION_SCHEMA.ROUTINES | Lists routines |
| Execute routine | CALL <dataset>.<routine>(test_params) | Returns expected result |
| Permissions check | bq show --format=prettyjson <dataset> | ACLs include Flume’s service account |
Every connection starts from the pre-engagement checklist and goes through the universal validation protocol before production sign-off.