Skip to main content
Dodaj swoje CV - zajmie to tylko kilka sekund

Oferty pracy: intern, data science

Sortuj według: -

Podobne wyszukiwania:

machine learning internship

Job Post Details

Ta oferta pracy wygasła w serwisie Indeed
Powody mogą być następujące: pracodawca nie przyjmuje aplikacji, nie prowadzi aktywnej rekrutacji lub przegląda otrzymane aplikacje.

Lead Data Engineer - job post

Qubit Labs
Poznań, wielkopolskieHybrydowo
Od 22 000 PLN za miesiąc - Pełny etat

Lokalizacja

Poznań, wielkopolskieHybrydowo

Pełny opis stanowiska

Our client is looking for a Lead Data Engineer who will help Financial Institutions combat money laundering and fraud by building high-quality, governed, and resilient data platforms that power advanced analytics and financial crime detection.
In this role, you will shape and evolve the Databricks + AWS lakehouse architecture, enabling investigators and product teams to uncover criminal patterns and act effectively.
You will work within the development team, focusing on designing modern data solutions using Databricks/Snowflake.

Requirements

Deep expertise in SQL and hands-on experience with Databricks, Snowflake, Python, and PySpark to deliver advanced data engineering solutions.

Strong background in designing scalable and reusable data models, pipelines, and frameworks using Hadoop, Apache NiFi, and modern Cloud Data Lake architectures.

Experience with orchestration tools: Airflow (DAGs, sensors, task groups), Databricks Workflows, AWS Step Functions.

Solid knowledge of the AWS data ecosystem: S3 structuring, IAM (least privilege), Glue Catalog, Lake Formation, VPC networking, encryption.

Proficient in CI/CD practices (Git branching, PR workflows, automated deployments, IaC with Terraform/CloudFormation).

Familiarity with governance and lineage tools (Unity Catalog, OpenLineage, Atlas, or similar), and understanding of audit/compliance requirements (PII/PCI, retention).

Strong production experience optimizing large-scale Spark/PySpark pipelines in Databricks (clusters, jobs, Delta Lake, Photon).

Nice-to-have: Python packaging, OpenTelemetry, Financial Crime domain experience.

Proven success working with stakeholders to translate business needs into reliable, high-impact cloud data solutions (Azure, AWS, cloud DWH).

Strong leadership and communication abilities to guide and mentor engineering teams and collaborate across functions.

Cost optimization skills (storage tiering, right-sizing, cache strategies, spot instances).

Comfortable leading design discussions, coaching engineers, and working with teams across data science, security, compliance, and product.

Practical mindset with focus on robustness, speed, automation, and clear documentation.

Experience supporting incident management (on-call, observability, reducing MTTR).

Responsibilities

Lead end-to-end design, development, optimization, and support of scalable Spark/PySpark pipelines in Databricks (batch & streaming).

Set and enforce standards for lakehouse & Medallion architecture (bronze/silver/gold), governance, data quality SLAs, lineage, and cost management.

Manage data ingestion using Apache NiFi, SFTP/FTPS, APIs, ensuring secure and reliable onboarding of internal/external datasets.

Architect compliant AWS data infrastructure (S3, IAM, KMS, Glue, Lake Formation, Lambda, Step Functions, CloudWatch, Secrets Manager, EC2/EKS).

Implement orchestration through Airflow, Databricks Workflows, and Step Functions, standardizing DAG best practices.

Drive data quality initiatives (expectations, anomaly detection, reconciliation, tests) and ensure reliability via SLIs/SLOs and effective alerting.

Enable metadata and lineage integration (Unity Catalog, Glue, OpenLineage) to support audit, impact analysis, and regulatory needs.

Lead CI/CD for data assets, including IaC, notebook/test automation, versioning, semantic tagging, and environment promotion.

Mentor engineers on distributed processing, partitioning, Delta Lake optimizations, caching, and cost–performance trade-offs.

Collaborate with data science, product, and compliance teams to translate analytical and detection requirements into robust data models and serving layers.

Perform code reviews (PySpark, SQL, infra templates) and lead architecture/design sessions.

Implement secure secrets management, key rotation, data masking/tokenization, and fine-grained access controls.

Support continuous improvement: handle backlog, sizing, delivery tracking, and stakeholder demos.

Participate in incident response, including root cause analysis, postmortems, and preventive engineering.

Hybrid role; office attendance per Mastercard policy.

Job Type: Full-time

Pay: From 22,000.00zł per month

Aplikuj łatwo na oferty pracyStwórz swoje CV