AWS Glue
Glue Data Catalog, Glue ETL jobs, Glue crawlers.
Access mode: Read-only
Required information
| Field | Details |
|---|---|
| AWS Account ID | Account containing Glue resources. |
| Region | AWS region. |
| IAM role | Cross-account role with Glue read permissions. |
| Scope | Glue Data Catalog (databases, tables), ETL job definitions, or both. |
Network considerations
AWS managed service. HTTPS to glue.<region>.amazonaws.com. No VPN needed.
VPC endpoint: If Flume runs in AWS and client wants private connectivity, create an interface VPC endpoint for Glue.
Credential and auth management
Preferred: Cross-account IAM role assumption. Role with glue:GetDatabases, glue:GetTables, glue:GetJobs, glue:GetCrawlers.
Glue Data Catalog permissions: May also need Lake Formation permissions if catalog is governed by Lake Formation.
No session concept: Each API call independently authenticated via SigV4.
Validation checks
| Check | Method | Expected result |
|---|---|---|
| Authentication | aws glue get-databases | Returns database list |
| Table access | aws glue get-tables --database-name <db> | Lists tables |
| Job listing | aws glue get-jobs | Lists ETL jobs |
| Job definition | aws glue get-job --job-name <name> | Returns job script location and config |
| Crawler listing | aws glue get-crawlers | Lists crawlers |
Every connection starts from the pre-engagement checklist and goes through the universal validation protocol before production sign-off.