Software & Engineering · Data & ML · 30-minute interview

Machine Learning Engineer interview questions and practice.

Builds, trains and deploys machine learning models into production systems, combining software engineering with data science.

No card for the taster. Full interviews are paid one at a time. Nothing renews.

Last reviewed

This page is still being written: no authored question bank for this competency family. The role is fully supported in the interview itself; only the published question bank is outstanding.

8 scored competencies30-minute voice interviewScored in about a minute after the call

What interviewers for Machine Learning Engineer actually ask

The question bank for this role is still being written. These are the first three competencies in the model the interview is scored against.

  1. Translates a business or product goal into a well-posed ML task with a metric that reflects real value, and knows when ML is not the right tool.

    Problem framing & success metrics
  2. Understands and validates the data before modelling, engineers features with awareness of leakage and drift, and reproduces data pipelines reliably.

    Data quality & feature engineering
  3. Selects and trains appropriate models, evaluates them honestly with proper baselines and error analysis, and can explain why a model made a prediction.

    Modelling & rigorous evaluation

What they are really assessing

Interviewers rarely score whether you seemed nice. They score against a model like this one, usually without telling you it exists. Each competency has a weak, adequate and strong shape, and the difference is almost always the level of specific detail you volunteer without being asked.

Problem framing & success metrics

Translates a business or product goal into a well-posed ML task with a metric that reflects real value, and knows when ML is not the right tool.

Weak
Starts from the model rather than the problem; success metric is accuracy regardless of class balance or cost of errors; cannot describe a case where they recommended not using ML.
Adequate
Frames the task correctly and picks a sensible offline metric with awareness of imbalance and error costs, but the link between offline metric and business outcome is asserted rather than validated.
Strong
Explains how they chose the target and metric from the decision it supports, how offline and online metrics diverged in a real project, what a baseline or heuristic achieved, and a case where they stopped an ML project.

Data quality & feature engineering

Understands and validates the data before modelling, engineers features with awareness of leakage and drift, and reproduces data pipelines reliably.

Weak
Takes the dataset as given; cannot describe a leakage bug, a label quality problem or how train/test splits respected time; features are whatever the library produced.
Adequate
Checks distributions and missingness, splits by time when relevant, and has caught one leakage or label problem, but data validation is manual and not repeatable.
Strong
Gives a concrete leakage or label bug they found, how it inflated results, and the automated validation added; explains feature choices with evidence and how they monitor feature drift in production.

Modelling & rigorous evaluation

Selects and trains appropriate models, evaluates them honestly with proper baselines and error analysis, and can explain why a model made a prediction.

Weak
Model choice is whatever was in the tutorial; evaluation is a single number on a random split; cannot describe error analysis or explain a failure mode of the model.
Adequate
Compares against a baseline, uses cross-validation or holdout appropriately, and has done error analysis, but hyperparameter and architecture decisions are not clearly justified.
Strong
Describes a real modelling decision with the baselines beaten, error analysis that changed the approach, calibration or fairness checks, and the experiment tracking used to make results reproducible.

Deploying & operating ML in production

Ships models as reliable services or batch jobs with monitoring for drift, latency and quality, and can retrain, roll back and explain incidents.

Weak
Experience ends at a notebook; cannot describe how a model was served, what was monitored, or what happened when performance degraded.
Adequate
Has deployed a model behind an API or batch job with basic latency and error monitoring, but drift detection and retraining are manual and there was no formal rollback path.
Strong
Gives a production incident (drift, skew between training and serving, latency spike), how it was detected, the rollback or retrain, and the monitoring, feature store or CI checks introduced afterwards.

LLM & generative AI systems

Builds applications on foundation models using prompting, retrieval, fine-tuning and evaluation, with attention to cost, latency, hallucination and safety.

Weak
Experience is calling an API with a prompt; no evaluation beyond eyeballing; cannot explain when retrieval or fine-tuning is appropriate or how to detect hallucination.
Adequate
Has built a RAG or agentic system with a basic eval set and guardrails, but evaluation is small and manual, and cost/latency trade-offs were not measured systematically.
Strong
Describes an LLM system with a versioned eval set, measured hallucination and quality, chunking and retrieval decisions backed by results, cost/latency numbers per request, and a failure caught by evals before release.

Experimentation & causal thinking

Validates that a model actually improves outcomes through A/B tests or other causal methods, and avoids fooling themselves with offline gains.

Weak
Assumes a better offline metric means a better product; cannot describe an A/B test they ran, its sample size reasoning, or a result that was noise.
Adequate
Has run an A/B test with a pre-registered metric and reasonable duration, but cannot discuss power, peeking, or what to do when offline and online results disagreed.
Strong
Explains an experiment with power calculation, guardrail metrics and a surprising result, how they diagnosed it, and a case where they refused to ship despite an offline improvement.

Communicating results & uncertainty

Explains model behaviour, limitations and confidence to non-specialists so they make good decisions, and is honest about what the model cannot do.

Weak
Presents a single accuracy number as the result; cannot explain confidence, failure modes or bias to a product manager; overstates what the model can do.
Adequate
Explains limitations and confidence intervals when asked and has pushed back on an unrealistic expectation, but tends to present results in technical terms.
Strong
Gives a case where they changed a business decision by explaining uncertainty or a failure mode clearly, including how they framed it for the audience and what safeguards they recommended.

Responsible AI, fairness & data privacy

Assesses bias, fairness and privacy risks in data and models, and applies appropriate safeguards and regulatory awareness (e.g. POPIA/GDPR consent and purpose limitation).

Weak
Treats fairness and privacy as someone else's job; has never checked model performance across groups or considered whether personal data was used lawfully.
Adequate
Has evaluated performance across demographic slices and anonymised data where required, but mitigation and documentation were ad hoc.
Strong
Describes a fairness or privacy issue they found (disparate error rates, unconsented data), how they mitigated it and documented it, and how they built checks into the pipeline.

Reading the questions is the easy half. Try answering three of them out loud, to someone who follows up.

Try 5 minutes free

What your 30 minutes covers

The same shape as a real first-round interview, pitched at mid-level Machine Learning Engineer and scored throughout.

0 to 7 min

Warm-up, then Motivation & fit

Build rapport, settle nerves, and get a short walk-through of your background. Why this role, why this employer, and what you are actually looking for.

7 to 16 min

Your experience

Two or three real situations from your CV in depth: context, what you did, what happened, what you would change.

Pitched at mid-level scope: owns a model or ML feature end to end from framing to production monitoring; accountable for its quality and its business metric.

16 to 25 min

Role-specific questions

The core competencies and domain knowledge for the role, with follow-ups on anything vague.

Drawn from this role's domain: choosing a target and evaluation metric for an imbalanced classification problem, detecting and preventing data leakage and training-serving skew and building a repeatable training pipeline with experiment tracking and versioned data, and the rest of the competency model.

25 to 30 min

Your questions, then Wrap-up

Your questions for the interviewer, and yes, they are assessed. Next steps and a clean finish.

What changes with seniority

The questions barely change between levels. What changes is the answer they will accept.

 JuniorMidSenior
Scope of ownershipOwns feature engineering, training and evaluation for a defined model with review; expected to validate data and write reproducible pipelines.Owns a model or ML feature end to end from framing to production monitoring; accountable for its quality and its business metric.Owns an ML domain or platform (recommendations, fraud, LLM features), its roadmap and technical standards; accountable for outcomes across multiple models.
Tolerance for ambiguityHandles a defined problem with some data gaps; escalates unclear targets or metrics rather than guessing.Turns a business problem into an ML formulation, chooses metrics, and challenges the ask when ML is not the answer.Defines what to build from business strategy; makes build-vs-buy and architecture decisions with incomplete evidence and owns them.
People leadershipNone formally.Mentors juniors, reviews experiments, may lead a small project.Technical lead for ML engineers and scientists, drives review standards, mentors across teams.
Who they deal withOwn team, data engineers, occasionally a product owner.Product managers, data engineers, platform teams, business stakeholders who consume the model output.Product and data leadership, engineering platform teams, legal/privacy, executives consuming results.

What your report would say

Every competency above scored from your own answers, the sentence that cost you quoted back, and your weakest answers rewritten the way a strong Machine Learning Engineer would have said them.

Sample report · Machine Learning Engineer
Mid-level · Mixed · 30:00
64of 100
Competencies, scored
Problem framing & success metrics4/5
Data quality & feature engineering3/5
Modelling & rigorous evaluation2/5
Deploying & operating ML in production3/5
LLM & generative AI systems4/5
What strong looks like: Modelling & rigorous evaluation
  • Describes a real modelling decision with the baselines beaten, error analysis that changed the approach, calibration or fairness checks, and the experiment tracking used to make results reproducible.

The format, not a result. Scores on your report come from what you actually said.

Is the AI interviewer realistic? See a full sample report

Related roles

Fail this interview here, not there.

Thirty minutes with a demanding Machine Learning Engineer interviewer now is the cheapest way to find out what you would have got wrong later.