
Senior Data Engineer (GCP, PySpark, Dataproc)
Description
We are looking for a Senior Data Engineer to join Implex and work on a long-term data platform project for a public-sector in the Middle East.
The role combines hands-on PySpark development with the configuration and troubleshooting of data integrations across GCP Dataproc, S3-compatible storage, PostgreSQL, and SAP HANA.
You will work with large-scale ETL pipelines, investigate Spark and infrastructure-level issues, and help ensure reliable and secure data exchange between multiple systems.
Requirements
Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices
Responsibilities
Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
Configure and troubleshoot Spark workloads running on GCP Dataproc.
Package and submit jobs, manage dependencies, and analyze driver and executor logs.
Identify and resolve Spark performance, memory, partitioning, and data processing issues.
Integrate Spark with S3-compatible storage using the Hadoop S3A connector.
Configure and troubleshoot MinIO connectivity, bucket access, TLS endpoints, signatures, and redirects.
Configure custom CA certificates and Java truststores for Spark drivers and executors.
Implement scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA.
Build reliable incremental data loads, retries, backfills, and data reconciliation processes.
Contribute to data modeling, schema evolution, documentation, and data quality practices.
Nice to Have
Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
KAFKA knowledge if we ever bring KAFKA into the architecture again
Tech Stack
Industries
Benefits
Implex
Outsource