No Relocation Assistance Offered
Job Number #170957 - Piscataway, New Jersey, United States
The Director of Data Engineering & Analytics leads the strategic direction and technical implementation of our data ecosystem. This is a Player-Coach role: you will lead a team and define the roadmap, but you will also roll up your sleeves to architect solutions, contribute code, and drive technical excellence.
You will bridge the gap between complex scientific data (lab experiments, formulation efficacy) and actionable business insights. You will own the full data lifecycle—from ingestion and pipeline orchestration to governance and final visualization—ensuring our R&D teams have the data they need to innovate faster.
1. Technical Leadership & Architecture (The "Player")
Hands-on Engineering: Architect and maintain robust data pipelines using Airflow and Python. Actively contribute to the codebase and perform code reviews to ensure high standards.
Infrastructure Management: Oversee container orchestration using Kubernetes and manage CI/CD workflows via GitHub Actions to ensure seamless deployment.
Data Modeling: Design scalable data models that can ingest and harmonize disparate data sources, including complex scientific data sets, and consumer feedback.
Tech Stack Evolution: Evaluate and implement modern tools (e.g., Ingestion tools, Cloud Data Warehouses, ML Ops, and AI Ops) to future-proof our capabilities.
2. Strategy & Team Management (The "Coach")
Strategic Roadmapping: Act as the Product Owner for the data platform. Partner with R&D, Supply Chain, and Quality leaders to define KPIs and build a roadmap that directly impacts product quality and innovation.
Team Mentorship: Lead, mentor, and grow a team of analytics engineers and data scientists. Foster a culture of engineering excellence, curiosity, and scientific rigor.
Data Governance: Establish and enforce data governance policies to ensure the integrity, reproducibility, and security of scientific data (essential for regulatory and quality compliance).
Translation: Translate complex technical challenges into clear insights for non-technical stakeholders, ensuring data is accessible and understandable across the organization.
Education: Bachelor’s degree in Computer Science, Data Science, Engineering, or a related scientific field.
Experience:
7+ years of experience in scientific and/or data-related roles, with a focus on product quality and R&D applications.
The "Stack": Proven expertise in:
Languages: Advanced Python and SQL skills.
Orchestration: Apache Airflow (or similar workflow management tools).
DevOps/Infra: Terraform, Kubernetes, Docker, and CI/CD (specifically GitHub Actions).
Version Control: Git flow best practices.
Master’s degree in Computer Science, Data Science, Engineering, or a related scientific field.
Experience with Cloud Platforms (GCP).
Familiarity with Modern Data Warehouses (Snowflake).
Experience implementing and maintaining Machine Learning models for formulation/optimization.
Background in Life Sciences, biotech, or consumer goods R&D.
Domain Knowledge: Experience working with scientific, R&D, clinical, or manufacturing data is highly preferred. You understand that "data integrity" in a science context means reproducibility and precision.
Soft Skills: Excellent ability to translate complex data into practical insights for a scientific audience. Strong communication and collaboration skills to work effectively with both technical and non-technical stakeholders.
Compensation and Benefits
Salary Range $175,000.00 - $200,000.00 USD