Data Analyst interview prep
A spaced-repetition deck of 265+ Data Analyst interview questions — organised by topic and difficulty, and scheduled for review based on your ratings. Preview a few cards below, then choose access to study the whole track on an Anki-style SM-2 schedule.
Eligible new monthly or yearly subscribers get 7 days free. Checkout confirms eligibility and charges.
Before you start
Start with basic SQL and arithmetic/statistics. Python and pandas modules require reading basic Python; practise queries and analyses on a small dataset alongside the cards.
First pass: Databases → SQL for Analytics. Begin with the foundational questions in these modules, then follow the outline. Revisit advanced and senior questions as the job requires; all cards remain available from the start.
Shared topics can appear in several tracks and keep one review history. The outline labels shared foundations and role-specific modules; choose your depth for the job you are preparing for.
What's covered
Every topic in this track, grouped the way you'd study it.
Databases
42 cardsShared foundations
SQL for Analytics
11 cardsRole-specific module
Statistics & Probability
11 cardsRole-specific module
A/B Testing & Experimentation
9 cardsRole-specific module
Product & Business Metrics
10 cardsRole-specific module
Data Visualization & Storytelling
8 cardsRole-specific module
Python
122 cardsShared foundations
Python for Analysts (pandas & matplotlib)
10 cardsRole-specific module
Analytical Case Studies
7 cardsRole-specific module
Behavioral
35 cardsShared foundations
Sample questions
A few cards from the deck — reveal each answer, then choose access to study the full set on a schedule.
What does SELECT do, and in what order do the parts of a query execute?
What does SELECT do, and in what order do the parts of a query execute?
Short answer: SELECT retrieves rows from tables. The logical execution order does NOT match the written order: first FROM, then WHERE, GROUP BY, HAVING, SELECT, DISTINCT, ORDER BY, LIMIT.
In depth:
Written (syntactic) order:
SELECT DISTINCT col1, agg(col2)
FROM t
WHERE cond
GROUP BY col1
HAVING agg_cond
ORDER BY col1
LIMIT 10 OFFSET 20;
Logical execution order:
1. FROM / JOIN -- which tables, how to join
2. WHERE -- row filter BEFORE grouping
3. GROUP BY -- grouping
4. HAVING -- group filter AFTER aggregation
5. SELECT -- evaluate expressions, aliases
6. DISTINCT -- remove duplicates
7. ORDER BY -- sorting
8. LIMIT / OFFSET -- slice
⚠️ Gotcha: an alias from SELECT cannot be used in WHERE (since WHERE runs before SELECT), but it can be used in ORDER BY and often in GROUP BY (depends on the DBMS). Example of the error:
SELECT salary * 12 AS annual FROM emp WHERE annual > 100000; -- ERROR
SELECT salary * 12 AS annual FROM emp ORDER BY annual; -- OK
What are window functions and how do they differ from GROUP BY?
What are window functions and how do they differ from GROUP BY?
Short answer: A window function computes an aggregate over a "window" of rows but does not collapse them — every source row survives and gets its own value. GROUP BY instead folds each group into a single row.
In depth:
- GROUP BY — N rows in a group → 1 result row. Detail is lost.
- Window function — N rows stay N rows, with an aggregate (sum, rank, average, lag) added alongside.
- Syntax —
func() OVER (PARTITION BY ... ORDER BY ...).PARTITION BYdefines the groups,ORDER BYsets order within the window (needed for running totals and ranks).
-- Running total of sales per user
SELECT
user_id,
order_date,
amount,
SUM(amount) OVER (
PARTITION BY user_id
ORDER BY order_date
) AS running_total
FROM orders;
⚠️ Common mistake: trying to filter on a window function's result in WHERE — windows are evaluated after WHERE/GROUP BY, so wrap the query in a CTE/subquery and filter on the outside.
How do mean, median, and mode differ, and why is the median better for skewed data (income, latency)?
How do mean, median, and mode differ, and why is the median better for skewed data (income, latency)?
Short answer: The mean is the sum divided by the count and is sensitive to outliers. The median is the middle value of the ordered data and is robust to outliers. The mode is the most frequent value. For skewed data the median better reflects the "typical" observation.
In depth:
- Mean — good for symmetric data and downstream math (variance, regression), but a single outlier shifts it.
- Median — the 50th percentile; splits the data in half and barely reacts to extremes.
- Mode — the only one of the three that works for categorical data ("most common plan").
| Metric | When to use | Outlier sensitivity |
|---|---|---|
| Mean | symmetric data, downstream calculations | high |
| Median | skewed data: income, latency, prices | low |
| Mode | categorical data, most frequent value | low |
⚠️ Common mistake: reporting the mean for income or latency. A long right tail drags the mean upward, so it no longer describes the typical user — use the median.
How does an A/B test work end to end? Walk through the key steps.
How does an A/B test work end to end? Walk through the key steps.
Short answer: An A/B test is a randomized experiment: users are split at random into control and treatment, shown different versions, and compared on a single pre-chosen primary metric under a fixed decision rule.
In depth:
- Hypothesis — what you change, which metric it should move, and why.
- Metric and MDE — one primary metric plus a minimum detectable effect (MDE).
- Randomization — a random, stable split into control/treatment.
- Sample size and duration — compute and fix them before launch.
- Run to N — run until the sample is reached, without peeking at interim results.
- Analysis — compare the metric, compute the p-value and the confidence interval of the lift.
- Decision — ship or roll back per the pre-defined rule.
Hypothesis → Randomize → Run to N → Analyze → Decision
│ │
control / treatment ship ✓ / rollback ✗
⚠️ Common mistake: changing the UI mid-test or stopping the moment it first looks 'significant' — both break the conclusions.
What makes a metric good, and how does a metric differ from a KPI?
What makes a metric good, and how does a metric differ from a KPI?
Short answer: A good metric is measurable, comparable, understandable and hard to game, and above all actionable: it can drive a decision. A KPI is the same kind of metric but tied to a goal and a target value for a period; a KPI is a subset of metrics, not a synonym.
In depth:
- Actionable, not vanity — a metric should change someone's decision. If the number goes up but nobody knows what to do, it's a vanity metric.
- Comparability — prefer rates and ratios (conversion, retention %) over raw counts: they compare across periods and segments.
- Hard to game — ask "how could this be inflated without creating value?". "Clicks" are easy to game; "activated users" are harder.
- A precise definition — one numerator, one denominator, a fixed window and filters, or two teams will count it differently.
| Property | Metric | KPI |
|---|---|---|
| What it is | any measurable number | a metric with a goal |
| Tied to a target | not necessarily | always (target) |
| How many | many | a chosen few |
| Example | time on page | activation rate ≥ 40% by Q3 |
⚠️ Common mistake: calling everything a KPI. KPIs are the 3–5 metrics a team is actually judged on; the rest are diagnostic metrics.
How do you pick the right chart type: bar, line, scatter, or pie?
How do you pick the right chart type: bar, line, scatter, or pie?
Short answer: The chart type is dictated not by taste but by the kind of relationship in the data: comparing categories — bars, change over time — a line, correlation between two variables — a scatter plot, share of a few parts in a whole — a pie. When torn between a pie and anything else, almost always pick bars.
In depth:
First state which relationship you are showing, and only then choose the visual:
| Relationship in data | Chart type | Why |
|---|---|---|
| Comparing categories | Bar | the eye compares length precisely; axis must start at 0 |
| Change over time | Line | a line emphasizes trend and continuity |
| Correlation of X and Y | Scatter | reveals the cloud, clusters and outliers |
| Parts of a whole | Pie / stacked | only for 2–4 parts, summing to 100% |
| Distribution | Histogram / boxplot | shows shape, spread, tails |
- Bars — the workhorse for categories; the Y axis starts at zero, otherwise the height difference lies.
- Line — only for an ordered continuous axis (time, days); never connect unrelated categories with a line.
- Scatter — for "are X and Y related"; add a trend line if needed.
⚠️ Common mistake: a pie chart with 8 slices — the shares become indistinguishable; replace it with horizontal bars sorted by value.
Ready to make it stick?
Explore a sample, choose a track, and build a review routine that fits your preparation.
Questions about this track
How should I prepare for a Data Analyst interview?
Study the concepts you'll be asked to explain, not just the ones you can code. RecallDeck's Data Analyst track gives you 265+ curated interview questions and resurfaces each one with an Anki-style SM-2 schedule based on your ratings — to practise recalling and explaining them. Combine reviews with coding and mock interviews.
What topics does the Data Analyst track cover?
The Data Analyst track is organised into the core areas Data Analyst interviews actually test, grouped by topic and by difficulty (Concept, Junior, Middle, Senior). You can preview the full outline and sample questions above before signing in.
Is spaced repetition effective for Data Analyst interview prep?
Yes. Active recall and spaced practice can help retain what you study. Grade yourself honestly and check your understanding with practical work. RecallDeck schedules each Data Analyst card using your ratings, so familiar cards return less often and difficult ones get more practice.
Can I try the Data Analyst track before paying?
Eligible new monthly or yearly subscribers can try the complete Data Analyst track and every feature for seven days. Checkout confirms trial eligibility and the amount due; returning subscribers may be charged immediately. Cancel online before the trial ends to avoid its first charge. Lifetime access has no trial.