AI Engineer interview questions and practice.
Designs and integrates AI capabilities such as large language models, agents and computer vision into products and workflows.
No card for the taster. Full interviews are paid one at a time. Nothing renews.
Last reviewed
This page is still being written: no authored question bank for this competency family. The role is fully supported in the interview itself; only the published question bank is outstanding.
What interviewers for AI Engineer actually ask
The question bank for this role is still being written. These are the first three competencies in the model the interview is scored against.
Translates a business or product goal into a well-posed ML task with a metric that reflects real value, and knows when ML is not the right tool.
Understands and validates the data before modelling, engineers features with awareness of leakage and drift, and reproduces data pipelines reliably.
Selects and trains appropriate models, evaluates them honestly with proper baselines and error analysis, and can explain why a model made a prediction.
What they are really assessing
Interviewers rarely score whether you seemed nice. They score against a model like this one, usually without telling you it exists. Each competency has a weak, adequate and strong shape, and the difference is almost always the level of specific detail you volunteer without being asked.
Problem framing & success metrics
Translates a business or product goal into a well-posed ML task with a metric that reflects real value, and knows when ML is not the right tool.
- Weak
- Starts from the model rather than the problem; success metric is accuracy regardless of class balance or cost of errors; cannot describe a case where they recommended not using ML.
- Adequate
- Frames the task correctly and picks a sensible offline metric with awareness of imbalance and error costs, but the link between offline metric and business outcome is asserted rather than validated.
- Strong
- Explains how they chose the target and metric from the decision it supports, how offline and online metrics diverged in a real project, what a baseline or heuristic achieved, and a case where they stopped an ML project.
Data quality & feature engineering
Understands and validates the data before modelling, engineers features with awareness of leakage and drift, and reproduces data pipelines reliably.
- Weak
- Takes the dataset as given; cannot describe a leakage bug, a label quality problem or how train/test splits respected time; features are whatever the library produced.
- Adequate
- Checks distributions and missingness, splits by time when relevant, and has caught one leakage or label problem, but data validation is manual and not repeatable.
- Strong
- Gives a concrete leakage or label bug they found, how it inflated results, and the automated validation added; explains feature choices with evidence and how they monitor feature drift in production.
Modelling & rigorous evaluation
Selects and trains appropriate models, evaluates them honestly with proper baselines and error analysis, and can explain why a model made a prediction.
- Weak
- Model choice is whatever was in the tutorial; evaluation is a single number on a random split; cannot describe error analysis or explain a failure mode of the model.
- Adequate
- Compares against a baseline, uses cross-validation or holdout appropriately, and has done error analysis, but hyperparameter and architecture decisions are not clearly justified.
- Strong
- Describes a real modelling decision with the baselines beaten, error analysis that changed the approach, calibration or fairness checks, and the experiment tracking used to make results reproducible.
Deploying & operating ML in production
Ships models as reliable services or batch jobs with monitoring for drift, latency and quality, and can retrain, roll back and explain incidents.
- Weak
- Experience ends at a notebook; cannot describe how a model was served, what was monitored, or what happened when performance degraded.
- Adequate
- Has deployed a model behind an API or batch job with basic latency and error monitoring, but drift detection and retraining are manual and there was no formal rollback path.
- Strong
- Gives a production incident (drift, skew between training and serving, latency spike), how it was detected, the rollback or retrain, and the monitoring, feature store or CI checks introduced afterwards.
LLM & generative AI systems
Builds applications on foundation models using prompting, retrieval, fine-tuning and evaluation, with attention to cost, latency, hallucination and safety.
- Weak
- Experience is calling an API with a prompt; no evaluation beyond eyeballing; cannot explain when retrieval or fine-tuning is appropriate or how to detect hallucination.
- Adequate
- Has built a RAG or agentic system with a basic eval set and guardrails, but evaluation is small and manual, and cost/latency trade-offs were not measured systematically.
- Strong
- Describes an LLM system with a versioned eval set, measured hallucination and quality, chunking and retrieval decisions backed by results, cost/latency numbers per request, and a failure caught by evals before release.
Experimentation & causal thinking
Validates that a model actually improves outcomes through A/B tests or other causal methods, and avoids fooling themselves with offline gains.
- Weak
- Assumes a better offline metric means a better product; cannot describe an A/B test they ran, its sample size reasoning, or a result that was noise.
- Adequate
- Has run an A/B test with a pre-registered metric and reasonable duration, but cannot discuss power, peeking, or what to do when offline and online results disagreed.
- Strong
- Explains an experiment with power calculation, guardrail metrics and a surprising result, how they diagnosed it, and a case where they refused to ship despite an offline improvement.
Communicating results & uncertainty
Explains model behaviour, limitations and confidence to non-specialists so they make good decisions, and is honest about what the model cannot do.
- Weak
- Presents a single accuracy number as the result; cannot explain confidence, failure modes or bias to a product manager; overstates what the model can do.
- Adequate
- Explains limitations and confidence intervals when asked and has pushed back on an unrealistic expectation, but tends to present results in technical terms.
- Strong
- Gives a case where they changed a business decision by explaining uncertainty or a failure mode clearly, including how they framed it for the audience and what safeguards they recommended.
Responsible AI, fairness & data privacy
Assesses bias, fairness and privacy risks in data and models, and applies appropriate safeguards and regulatory awareness (e.g. POPIA/GDPR consent and purpose limitation).
- Weak
- Treats fairness and privacy as someone else's job; has never checked model performance across groups or considered whether personal data was used lawfully.
- Adequate
- Has evaluated performance across demographic slices and anonymised data where required, but mitigation and documentation were ad hoc.
- Strong
- Describes a fairness or privacy issue they found (disparate error rates, unconsented data), how they mitigated it and documented it, and how they built checks into the pipeline.
Reading the questions is the easy half. Try answering three of them out loud, to someone who follows up.
Try 5 minutes freeWhat your 30 minutes covers
The same shape as a real first-round interview, pitched at mid-level AI Engineer and scored throughout.
Warm-up, then Motivation & fit
Build rapport, settle nerves, and get a short walk-through of your background. Why this role, why this employer, and what you are actually looking for.
Your experience
Two or three real situations from your CV in depth: context, what you did, what happened, what you would change.
Pitched at mid-level scope: owns a model or ML feature end to end from framing to production monitoring; accountable for its quality and its business metric.
Role-specific questions
The core competencies and domain knowledge for the role, with follow-ups on anything vague.
Drawn from this role's domain: choosing a target and evaluation metric for an imbalanced classification problem, detecting and preventing data leakage and training-serving skew and building a repeatable training pipeline with experiment tracking and versioned data, and the rest of the competency model.
Your questions, then Wrap-up
Your questions for the interviewer, and yes, they are assessed. Next steps and a clean finish.
What changes with seniority
The questions barely change between levels. What changes is the answer they will accept.
| Junior | Mid | Senior | |
|---|---|---|---|
| Scope of ownership | Owns feature engineering, training and evaluation for a defined model with review; expected to validate data and write reproducible pipelines. | Owns a model or ML feature end to end from framing to production monitoring; accountable for its quality and its business metric. | Owns an ML domain or platform (recommendations, fraud, LLM features), its roadmap and technical standards; accountable for outcomes across multiple models. |
| Tolerance for ambiguity | Handles a defined problem with some data gaps; escalates unclear targets or metrics rather than guessing. | Turns a business problem into an ML formulation, chooses metrics, and challenges the ask when ML is not the answer. | Defines what to build from business strategy; makes build-vs-buy and architecture decisions with incomplete evidence and owns them. |
| People leadership | None formally. | Mentors juniors, reviews experiments, may lead a small project. | Technical lead for ML engineers and scientists, drives review standards, mentors across teams. |
| Who they deal with | Own team, data engineers, occasionally a product owner. | Product managers, data engineers, platform teams, business stakeholders who consume the model output. | Product and data leadership, engineering platform teams, legal/privacy, executives consuming results. |
What your report would say
Every competency above scored from your own answers, the sentence that cost you quoted back, and your weakest answers rewritten the way a strong AI Engineer would have said them.
- Describes a real modelling decision with the baselines beaten, error analysis that changed the approach, calibration or fairness checks, and the experiment tracking used to make results reproducible.
The format, not a result. Scores on your report come from what you actually said.
Is the AI interviewer realistic? See a full sample report