Job Description…
We are seeking a highly experienced Senior Data Engineer with 10+ years of hands-on data engineering experience, including ownership of production data pipelines and at least one data platform or system managed end-to-end.
The ideal candidate will have strong expertise in modern data engineering technologies, including dbt, Apache Airflow, Python, SQL, Apache Spark/PySpark, Dremio, Apache Iceberg, S3-compatible storage, NoSQL databases, Kubernetes/OpenShift, GitLab CI/CD, and Docker.
The role requires extensive experience designing, developing, deploying, monitoring, and troubleshooting scalable production data pipelines and lakehouse architectures.
Must-Have Skills & Experience
Data Transformation & Modeling
- dbt (data build tool): models, tests, macros, incremental models, and project structure.
- Dimensional modeling and star schema.
- Medallion architecture: Raw/Curated/Presentation or Bronze/Silver/Gold.
- Data warehouse and lakehouse architecture.
Orchestration & Production Pipelines
- Apache Airflow: DAG design, dependencies, triggers, sensors, scheduling, parameterized runs, and failure handling.
- Production workflow orchestration, monitoring, and troubleshooting.
Data Ingestion & ETL/ELT
- Batch ingestion and ETL/ELT pipeline development.
- Full loads, incremental loads, upsert/merge, and Change Data Capture (CDC).
- Multi-source ingestion from SFTP, REST APIs, relational databases, and files.
- Data validation, data quality, and schema evolution.
Lakehouse, Query Engines & Storage
- Dremio or equivalent technologies such as Trino, Presto, Starburst, or Athena.
- Apache Iceberg or equivalent open table formats such as Delta Lake or Apache Hudi.
- AWS S3 or S3-compatible object storage including MinIO or Dell ECS.
- Partitioning, table layout, and query performance optimization.
Programming & Databases
- Strong Python skills for pipeline development, automation, and scripting.
- Advanced SQL, query optimization, and performance tuning.
- PostgreSQL and/or SQL Server.
NoSQL Databases
Hands-on data engineering experience with at least two of the following NoSQL database families:
- Redis or similar key-value/in-memory stores.
- MongoDB, Couchbase, or DocumentDB.
- Cassandra, ScyllaDB, or HBase.
- Elasticsearch or OpenSearch.
Candidates should have experience ingesting NoSQL data into a lakehouse, handling CDC/change streams, schema drift, semi-structured data, nested JSON, consistency models, TTLs, authentication, and performance considerations.
Data Formats & Semi-Structured Data
- Apache Parquet is required.
- Experience with ORC and Avro.
- JSON, JSON Lines, CSV/delimited files, and XML.
- Compression codecs including Snappy, gzip, and zstd.
- Nested and evolving schemas, schema inference, drift detection, and schema evolution.
Distributed Data Processing
- Apache Spark/PySpark or equivalent distributed processing technology.
- DataFrame API, joins, partitioning, skew handling, and writing Iceberg/Parquet.
- Understanding of parallelism, shuffles, memory optimization, and distributed processing costs.
Linux, Scripting & Connectivity
- Bash/shell scripting and Linux command line.
- SFTP/SSH, key management, automated transfers, archiving, and retries.
- TLS certificates and secure connectivity.
- Troubleshooting DNS, firewalls, proxies, timeouts, and Kubernetes/OpenShift networking.
- YAML and configuration management.
Pipeline Reliability
- Idempotent pipeline design and safe re-runs.
- Backfills, historical loads, reprocessing, and re-ingestion.
- Watermarking, late-arriving data, and deduplication.
- UTC/local time zone and scheduling considerations.
Security & Governance
- Kubernetes/OpenShift secrets and environment variables.
- PII awareness, data classification, and access controls.
- Audit logging and run-history frameworks.
- Data retention, archiving, and pruning policies.
Testing & Engineering Practices
- Python unit and integration testing using pytest.
- Mocking external systems and test fixture design.
- dbt schema and custom tests.
- Code review and linting using SQLFluff, ruff, or equivalent.
- Dependency and lockfile management.
APIs & Integrations
- REST API development and consumption using FastAPI, Flask, or equivalent.
- Pagination, rate limits, retries, and idempotency.
- API keys and OAuth2 client credentials.
- Webhooks, event-driven triggers, and integration platforms (iPaaS).
Platform & DevOps
- Kubernetes and/or OpenShift.
- Deployments, ConfigMaps, Secrets, volumes, and service accounts.
- GitLab repositories and GitLab CI/CD.
- Git version control and Dev/QA/Production environment promotion.
- Docker and containerized deployments.
- Change and release management.
Production Operations
- Production support and incident troubleshooting.
- Root-cause analysis across Airflow, dbt, and Kubernetes/OpenShift.
- Pipeline monitoring and alerting.
- Runbook authoring and runbook-driven operations.
Ways of Working
- Strong technical documentation skills, including READMEs, runbooks, and architecture diagrams.
- Experience using AI-assisted development tools such as Claude Code, GitHub Copilot, Cursor, or similar tools for development, debugging, and documentation.
Nice-to-Have Skills
- Nessie or other data catalog/versioning technologies.
- Advanced Apache Iceberg or Delta Lake knowledge.
- Apache Kafka or equivalent event-streaming technologies.
- Streaming ingestion patterns.
- Metadata management and data lineage.
- Data governance and classification.
- Experience working with regulated or government data.
- Experience taking over or stabilizing vendor-built data platforms.
- Arabic working proficiency.