S3 / GCS / Azure Blob Storage
Object storage: flat files, CSVs, Parquet, exports, and the data lake layer.
Access mode: Read-only (typically)
Required information
| Field | Details |
|---|---|
| Bucket/Container name(s) | S3 bucket, GCS bucket, or Azure Blob container names. |
| Prefix/path scope | Specific prefixes if not the entire bucket (e.g., data/exports/). |
| Cloud credentials | AWS: IAM role ARN. GCP: service account. Azure: service principal or SAS token. |
| File formats | What to expect: CSV, Parquet, JSON, Avro, ORC, etc. Affects parsing config. |
Network considerations
S3: HTTPS to s3.<region>.amazonaws.com. VPC endpoint (Gateway) keeps traffic on AWS backbone.
GCS: HTTPS to storage.googleapis.com.
Azure Blob: HTTPS to <account>.blob.core.windows.net. Private Endpoint available.
Bucket policies: Ensure Flume’s identity is allowed in bucket/container-level access policies.
Credential and auth management
Preferred: IAM role assumption (AWS). Cross-account role with s3:GetObject, s3:ListBucket.
Preferred: GCP service account. Storage Object Viewer role.
Preferred: Azure service principal. Storage Blob Data Reader role.
Azure SAS tokens: Acceptable for time-limited access. Set appropriate expiry and IP restrictions.
SSE/KMS: If objects are encrypted with CMK, Flume’s identity needs kms:Decrypt permission.
Validation checks
| Check | Method | Expected result |
|---|---|---|
| S3 access | aws s3 ls s3://<bucket>/<prefix>/ | Lists objects |
| GCS access | gsutil ls gs://<bucket>/<prefix>/ | Lists objects |
| Azure access | az storage blob list --container-name <c> --prefix <p> | Lists blobs |
| Read test | Download a sample file and parse | File content readable |
| Permission check | Attempt write (should fail for read-only) | Access denied confirms read-only |
Every connection starts from the pre-engagement checklist and goes through the universal validation protocol before production sign-off.