BI - Data Scientist
Cooper Standard · 9 hours ago
Job Description:
Role Overview
We are seeking a highly skilled Data Scientist with deep expertise in Databricks, Machine Learning, and AI-driven analytics to join our growing data organization. In this role, you will design, build, and deploy scalable models and data workflows that power enterprise insights, automation, and decision-making. You will work across engineering, analytics, and business teams to transform raw data into intelligent, production-ready solutions.
Key Responsibilities
- Develop, train, and deploy machine learning models using Databricks notebooks, MLflow, and the Lakehouse architecture.
- Build scalable ETL/ELT pipelines leveraging Delta Lake, PySpark, and Databricks workflows.
- Implement AI/LLM-based solutions, including retrieval-augmented generation (RAG), vector search, and enterprise agent workflows.
- Partner with data engineering to optimize datasets for analytics, modeling, and real-time inference.
- Conduct exploratory data analysis (EDA), feature engineering, and statistical modeling to uncover actionable insights.
- Use MLflow for experiment tracking, model versioning, and lifecycle management.
- Collaborate with business stakeholders to translate ambiguous problems into measurable, data-driven solutions.
- Deploy models into production using Databricks Model Serving, serverless compute, or API endpoints.
- Ensure governance, security, and compliance using Unity Catalog and enterprise data standards.
- Continuously evaluate new AI/ML technologies and recommend improvements to the platform and modeling strategy.
- Implement and uphold enterprise data governance standards, ensuring models and pipelines comply with regulatory, privacy, and audit requirements.
- Use Unity Catalog to manage secure, centralized governance for data, ML models, notebooks, and AI assets.
- Design and enforce Role-Based Access Control (RBAC) to ensure users only access data and models appropriate for their job functions.
- Apply Attribute-Based Access Control (ABAC) for fine-grained, dynamic access decisions based on user attributes (e.g., department, region, clearance level) and data attributes (e.g., sensitivity, classification).
Required Qualifications
- Bachelor’s or Master’s degree in Data Science, Computer Science, Statistics, or related field.
- 3–7+ years of experience building machine learning models in Python (Pandas, Scikit-learn, PySpark, TensorFlow, or PyTorch).
- Hands-on experience with Databricks, including notebooks, Delta Lake, MLflow, and Databricks SQL.
- Strong understanding of Lakehouse architecture, distributed computing, and scalable data processing.
- Experience deploying ML models into production environments.
- Proficiency in SQL and Python for data manipulation and analysis.
- Familiarity with LLMs, embeddings, vector databases, or AI agent frameworks.
- Ability to communicate complex technical concepts to non-technical stakeholders.
Preferred Qualifications
- Experience with Databricks Model Serving, Vector Search, or serverless warehouses.
- Background in NLP, deep learning, or generative AI.
- Experience integrating Databricks with SAP, Snowflake, or enterprise BI tools.
- Knowledge of MLOps best practices and CI/CD pipelines.
- Experience with cloud platforms (Azure, AWS, or GCP).
- Experience with SAP BDC, BDC Connect
What You’ll Bring
- A passion for solving complex problems with data and AI.
- Curiosity, creativity, and a strong desire to innovate.
- Ability to work in fast-paced, cross-functional environments.
- A mindset for building scalable, secure, and production-grade solutions.
Position Type:
Regular
Additional Locations:
Additional Information:
Remote Status:
Remote