Junior Data Analyst interview questions and practice.
Supports a data team by preparing data, running standard reports and learning to analyse and visualise results. An interviewer hiring a Junior Data Analyst is not testing whether you know what the job is. They are trying to establish whether the number you hand over would survive being questioned, and whether anyone would actually do something differently because of it.
No card for the taster. Full interviews are paid one at a time. Nothing renews.
Last reviewed
What interviewers for Junior Data Analyst actually ask
Three questions from the bank below, each scored against one competency. The follow-up is what separates a prepared answer from a memorised one.
Someone asks you for "a report on churn". What do you ask them first?
Tell me about an analysis where the question you were given was the wrong question.
Walk me through how you would find, in SQL, the customers who bought in January but not in February.
What they are really assessing
That gets scored against 7 competencies: framing the business question, sql & data wrangling, statistical reasoning & avoiding false conclusions, visualisation & data storytelling, data quality & validation discipline, stakeholder partnership & managing requests and ownership of impact & follow-through. Each one is assessed from the specifics in your answers, which is why "we improved the process" scores lower than a sentence with a number, a date and a decision in it.
At junior level they are testing whether you can be trusted with well-defined work and will ask for help before you break something. Expect them to push hardest on sQL competence including joins and window functions, a data error they caught, and explaining a finding in plain language.
Framing the business question
Turns a vague request into a precise, answerable question with a defined metric, population and time window, and checks the question is worth answering before pulling data.
- Weak
- Takes the request literally ('they asked for sales by region so I made a chart'); cannot say what decision the analysis was for or what metric definition was used.
- Adequate
- Clarifies the metric and scope with the requester and states the decision it supports, but assumptions are not written down and the analysis scope drifts.
- Strong
- Describes a real request they reframed: the clarifying questions asked, the metric definition agreed in writing, the decision it informed, and a case where they pushed back because the question would not change any decision.
SQL & data wrangling
Extracts, joins, cleans and reshapes data correctly using SQL and a scripting tool, and can spot when a join or filter has silently produced wrong numbers.
- Weak
- Writes basic SELECTs and relies on others for joins or window functions; cannot describe catching a fan-out join or duplicate rows; cleaning is done manually in a spreadsheet.
- Adequate
- Comfortable with joins, aggregations, CTEs and window functions and cleans data reproducibly in Python/R, but validation is ad hoc and they have been caught by a silent data error.
- Strong
- Gives a specific case where they caught a wrong result (fan-out, null handling, time zone, late-arriving data) by reconciling against a known total, and describes the checks they now build into every query.
Statistical reasoning & avoiding false conclusions
Applies the right level of statistical rigour: distinguishes noise from signal, correlation from causation, and knows when a sample size or comparison is misleading.
- Weak
- Reports any difference as a finding; cannot explain confidence intervals, seasonality or selection bias; has never questioned whether a trend was noise.
- Adequate
- Uses significance tests and seasonally adjusted comparisons and knows correlation is not causation, but struggles to explain p-values, power or a confounder in plain terms.
- Strong
- Gives a case where they stopped a wrong conclusion (small sample, Simpson's paradox, survivorship bias, regression to the mean), explains the reasoning in plain language and what they did to get a defensible answer.
Visualisation & data storytelling
Presents findings so a decision-maker grasps the point in seconds: the right chart, a clear headline, honest scales and the recommended action.
- Weak
- Dashboards and slides are data dumps; charts are chosen by default; cannot describe a presentation that led to a decision or feedback they received on clarity.
- Adequate
- Chooses appropriate charts, leads with a headline and keeps scales honest, but the narrative is descriptive rather than pointing to a decision or recommendation.
- Strong
- Describes a specific presentation: the one-line finding, the chart chosen over alternatives and why, how they handled a question that challenged the data, and the decision that was taken as a result.
Data quality & validation discipline
Checks data before trusting it, documents definitions and caveats, and reconciles numbers against known sources so results are right the first time.
- Weak
- Assumes the source is correct; cannot describe a data quality problem they found or a reconciliation they did; caveats are absent from reports.
- Adequate
- Profiles data for nulls, duplicates and ranges and reconciles totals to finance or another source, but checks are manual and not repeated when data refreshes.
- Strong
- Describes a data quality issue found (definition change, missing partition, duplicate loads), its business impact, how they fixed the numbers, and the automated check or documentation added so it does not recur.
Stakeholder partnership & managing requests
Works with business stakeholders as a partner: manages a queue of requests by value, says no or not yet with reasons, and builds trust that numbers are right.
- Weak
- Takes every request in order received; cannot describe declining or reprioritising a request, or dealing with a stakeholder who did not like the result.
- Adequate
- Prioritises requests by impact with their manager and has delivered an unwelcome result, but the stakeholder relationship was strained and they cannot say how they rebuilt it.
- Strong
- Gives an example of reprioritising a request with the stakeholder's agreement based on decision value, and of presenting a finding the stakeholder did not want, how they handled the pushback and what happened next.
Ownership of impact & follow-through
Follows analyses through to whether the decision was made and worked, and proactively finds problems in the data or the business rather than waiting for requests.
- Weak
- Work ends when the report is sent; cannot say whether any analysis changed an outcome; no example of an insight found without being asked.
- Adequate
- Tracks whether recommendations were adopted and has raised an unrequested insight, but cannot quantify the effect of their analysis on the business.
- Strong
- Gives an analysis with a measured business outcome (revenue, cost, churn) and an insight found proactively from routine monitoring that led to action, including how they followed up.
10 questions you should expect
What a strong answer contains, not a model answer to memorise. A memorised answer falls apart on the first follow-up, and there is always a follow-up.
Someone asks you for "a report on churn". What do you ask them first?
Scored against: Framing the business questionA strong answer contains: What decision this is for, what they would do differently depending on the answer, how churn is defined here and over what window, and what already exists. Producing the report without asking is the failure mode being tested.
And then they askThey say they just want to see the data. What now?
Tell me about an analysis where the question you were given was the wrong question.
Scored against: Framing the business questionA strong answer contains: The stated question, what they worked out the real one was, how they raised it without being obstructive, and what the analysis ended up being. This is the single strongest story an analyst can bring.
And then they askHow did you raise it without sounding like you were refusing the work?
Walk me through how you would find, in SQL, the customers who bought in January but not in February.
Scored against: SQL & data wranglingA strong answer contains: A correct approach said out loud (a left join with a null check, a NOT EXISTS, or an anti-join), plus the questions they would ask first: what counts as a purchase, which date, what about refunds, what time zone.
And then they askHow would your answer change if the orders table had fifty million rows?
What is the messiest data you have had to work with?
Scored against: SQL & data wranglingA strong answer contains: The specific problems (duplicated records, inconsistent keys, free-text categories, changing definitions over time), what they did about each, and what they documented so the next person did not repeat it.
And then they askWhat did you decide to exclude, and how did you justify it?
You have a result. How do you check it before you send it?
Scored against: Data quality & validation disciplineA strong answer contains: Actual checks: does the total reconcile to a known source, does the row count make sense, spot-check a handful of records by hand, compare to last period, ask whether the direction is plausible. Plus a case where a check caught something.
And then they askTell me about a time you sent out a number that was wrong. What happened?
A metric jumped 20% last week. Is that real?
Scored against: Statistical reasoning & avoiding false conclusionsA strong answer contains: Check the pipeline and definitions before believing it, then look at the base rate and variance, seasonality, a segment or campaign driving it, and whether the population changed. Correlation-versus-cause named without prompting.
And then they askHow would you tell a genuine change from noise?
How do you decide whether a difference between two groups matters?
Scored against: Statistical reasoning & avoiding false conclusionsA strong answer contains: Effect size as well as significance, the sample they had, what practical difference it would make to the decision, and honesty about the limits of an observational comparison.
And then they askThe p-value is 0.04 and the effect is tiny. What do you tell the stakeholder?
How do you present an analysis to people who will not read the appendix?
Scored against: Visualisation & data storytellingA strong answer contains: The answer first, then the two charts that support it, then the caveats, not a chronological account of the work. Includes a chart they simplified or removed because it was not doing anything.
And then they askWhat do you cut when you have five minutes instead of thirty?
A stakeholder does not like the answer and asks you to re-run it differently.
Scored against: Stakeholder partnership & managing requestsA strong answer contains: Distinguish a legitimate methodological challenge from a request for a different conclusion. Re-run it if the challenge is sound, hold the line if it is not, show both if it is genuinely ambiguous, and never quietly change a filter.
And then they askHow do you tell the difference between the two?
What is an analysis you did that actually changed something?
Scored against: Ownership of impact & follow-throughA strong answer contains: What was decided, by whom, and what happened afterwards. Analysts who can name the decision and the follow-up are rare; those who list dashboards built are not.
And then they askDid anyone check afterwards whether the decision worked?
Reading the questions is the easy half. Try answering three of them out loud, to someone who follows up.
Try 5 minutes freeWhat your 30 minutes covers
The same shape as a real first-round interview, pitched at junior Junior Data Analyst and scored throughout.
Warm-up, then Motivation & fit
Build rapport, settle nerves, and get a short walk-through of your background. Why this role, why this employer, and what you are actually looking for.
Your experience
Two or three real situations from your CV in depth: context, what you did, what happened, what you would change.
Pitched at junior scope: owns recurring reports and defined analyses end to end with review of conclusions.
Role-specific questions
The core competencies and domain knowledge for the role, with follow-ups on anything vague.
Drawn from this role's domain: reframing a vague stakeholder request into a measurable question, writing SQL with joins, window functions and CTEs and validating the result and catching a fan-out join, duplicate rows or null handling error, and the rest of the competency model.
Your questions, then Wrap-up
Your questions for the interviewer, and yes, they are assessed. Next steps and a clean finish.
What changes with seniority
The questions barely change between levels. What changes is the answer they will accept.
| Junior | Mid | Senior | |
|---|---|---|---|
| Scope of ownership | Owns recurring reports and defined analyses end to end with review of conclusions. | Owns analytics for a business area: request intake, analysis, dashboards and metric definitions; accountable for numbers being right. | Owns analytics strategy and metric governance for a department; leads complex cross-functional analyses; accountable for analytical quality. |
| Tolerance for ambiguity | Clarifies requests with the requester and fills small gaps sensibly. | Reframes vague requests into answerable questions and pushes back on low-value work. | Defines what questions the business should be asking; comfortable with open-ended problems and incomplete data. |
| People leadership | No formal leadership. | Mentors juniors, reviews their queries and reports. | Leads analysts technically, sets standards, may have direct reports. |
| Who they deal with | Own team, line manager, a few internal requesters. | Business managers, product or marketing teams, finance, data engineers. | Heads of department, executives, data engineering and BI leadership. |
Where candidates lose this interview
Answering the question as asked
The most valuable thing an analyst does is work out what the requester actually needs before writing any SQL. Candidates who go straight to the query (in the interview and, by implication, at work) are scored as report writers rather than analysts.
Numbers with no validation story
Interviewers ask how you check a result specifically because everyone has sent out a wrong one. If you cannot describe a reconciliation, a spot check or a sanity comparison, they assume nothing gets checked.
Charts described, decisions absent
"I built a dashboard the team uses daily" says nothing about whether anything changed. Have at least one story that ends with a decision, a person who made it and what happened next.
Statistical vocabulary without judgement
Saying "statistically significant" without effect size, sample or practical relevance is a common way to look junior. So is treating any correlation as a finding. Interviewers listen for the caveat you volunteer unprompted.
Quietly changing the analysis when challenged
Being pushed to re-run something until it says what a stakeholder wanted is the defining integrity test in this role. The strong answer distinguishes a real methodological objection from pressure, and says what it does with each.
What your report would say
Every competency above scored from your own answers, the sentence that cost you quoted back, and your weakest answers rewritten the way a strong Junior Data Analyst would have said them.
Someone asks you for "a report on churn". What do you ask them first?
- What decision this is for, what they would do differently depending on the answer, how churn is defined here and over what window, and what already exists. Producing the report without asking is the failure mode being tested
The format, not a result. Scores on your report come from what you actually said.
Is the AI interviewer realistic? See a full sample report