Interview track

Data Analyst interview prep

A spaced-repetition deck of 265+ Data Analyst interview questions — organised by topic and difficulty, and scheduled for review based on your ratings. Preview a few cards below, then choose access to study the whole track on an Anki-style SM-2 schedule.

265 cards13 topics
See access options

Eligible new monthly or yearly subscribers get 7 days free. Checkout confirms eligibility and charges.

Before you start

Start with basic SQL and arithmetic/statistics. Python and pandas modules require reading basic Python; practise queries and analyses on a small dataset alongside the cards.

First pass: Databases → SQL for Analytics. Begin with the foundational questions in these modules, then follow the outline. Revisit advanced and senior questions as the job requires; all cards remain available from the start.

Shared topics can appear in several tracks and keep one review history. The outline labels shared foundations and role-specific modules; choose your depth for the job you are preparing for.

What's covered

Every topic in this track, grouped the way you'd study it.

Databases

42 cards

Shared foundations

SQL Fundamentals

SQL for Analytics

11 cards

Role-specific module

SQL for Analytics

Statistics & Probability

11 cards

Role-specific module

Statistics

A/B Testing & Experimentation

9 cards

Role-specific module

Experimentation

Product & Business Metrics

10 cards

Role-specific module

Metrics

Data Visualization & Storytelling

8 cards

Role-specific module

Visualization

Python

122 cards

Shared foundations

Core LanguageData Model & InternalsConcurrency & AsyncStdlib, Typing & Testing

Python for Analysts (pandas & matplotlib)

10 cards

Role-specific module

Python for Analysts

Analytical Case Studies

7 cards

Role-specific module

Case Studies

Behavioral

35 cards

Shared foundations

Behavioral

Sample questions

A few cards from the deck — reveal each answer, then choose access to study the full set on a schedule.

What does SELECT do, and in what order do the parts of a query execute?

Short answer: SELECT retrieves rows from tables. The logical execution order does NOT match the written order: first FROM, then WHERE, GROUP BY, HAVING, SELECT, DISTINCT, ORDER BY, LIMIT.

In depth:

Written (syntactic) order:

SELECT   DISTINCT col1, agg(col2)
FROM     t
WHERE    cond
GROUP BY col1
HAVING   agg_cond
ORDER BY col1
LIMIT    10 OFFSET 20;

Logical execution order:

1. FROM / JOIN      -- which tables, how to join
2. WHERE            -- row filter BEFORE grouping
3. GROUP BY         -- grouping
4. HAVING           -- group filter AFTER aggregation
5. SELECT           -- evaluate expressions, aliases
6. DISTINCT         -- remove duplicates
7. ORDER BY         -- sorting
8. LIMIT / OFFSET   -- slice

⚠️ Gotcha: an alias from SELECT cannot be used in WHERE (since WHERE runs before SELECT), but it can be used in ORDER BY and often in GROUP BY (depends on the DBMS). Example of the error:

SELECT salary * 12 AS annual FROM emp WHERE annual > 100000; -- ERROR
SELECT salary * 12 AS annual FROM emp ORDER BY annual;        -- OK

What are window functions and how do they differ from GROUP BY?

Short answer: A window function computes an aggregate over a "window" of rows but does not collapse them — every source row survives and gets its own value. GROUP BY instead folds each group into a single row.

In depth:

  1. GROUP BY — N rows in a group → 1 result row. Detail is lost.
  2. Window function — N rows stay N rows, with an aggregate (sum, rank, average, lag) added alongside.
  3. Syntax — func() OVER (PARTITION BY ... ORDER BY ...). PARTITION BY defines the groups, ORDER BY sets order within the window (needed for running totals and ranks).
-- Running total of sales per user
SELECT
  user_id,
  order_date,
  amount,
  SUM(amount) OVER (
    PARTITION BY user_id
    ORDER BY order_date
  ) AS running_total
FROM orders;

⚠️ Common mistake: trying to filter on a window function's result in WHERE — windows are evaluated after WHERE/GROUP BY, so wrap the query in a CTE/subquery and filter on the outside.

How do mean, median, and mode differ, and why is the median better for skewed data (income, latency)?

Short answer: The mean is the sum divided by the count and is sensitive to outliers. The median is the middle value of the ordered data and is robust to outliers. The mode is the most frequent value. For skewed data the median better reflects the "typical" observation.

In depth:

  1. Mean — good for symmetric data and downstream math (variance, regression), but a single outlier shifts it.
  2. Median — the 50th percentile; splits the data in half and barely reacts to extremes.
  3. Mode — the only one of the three that works for categorical data ("most common plan").
Metric When to use Outlier sensitivity
Mean symmetric data, downstream calculations high
Median skewed data: income, latency, prices low
Mode categorical data, most frequent value low

⚠️ Common mistake: reporting the mean for income or latency. A long right tail drags the mean upward, so it no longer describes the typical user — use the median.

How does an A/B test work end to end? Walk through the key steps.

Short answer: An A/B test is a randomized experiment: users are split at random into control and treatment, shown different versions, and compared on a single pre-chosen primary metric under a fixed decision rule.

In depth:

  1. Hypothesis — what you change, which metric it should move, and why.
  2. Metric and MDE — one primary metric plus a minimum detectable effect (MDE).
  3. Randomization — a random, stable split into control/treatment.
  4. Sample size and duration — compute and fix them before launch.
  5. Run to N — run until the sample is reached, without peeking at interim results.
  6. Analysis — compare the metric, compute the p-value and the confidence interval of the lift.
  7. Decision — ship or roll back per the pre-defined rule.
Hypothesis → Randomize → Run to N → Analyze → Decision
              │                                  │
       control / treatment              ship ✓ / rollback ✗

⚠️ Common mistake: changing the UI mid-test or stopping the moment it first looks 'significant' — both break the conclusions.

What makes a metric good, and how does a metric differ from a KPI?

Short answer: A good metric is measurable, comparable, understandable and hard to game, and above all actionable: it can drive a decision. A KPI is the same kind of metric but tied to a goal and a target value for a period; a KPI is a subset of metrics, not a synonym.

In depth:

  1. Actionable, not vanity — a metric should change someone's decision. If the number goes up but nobody knows what to do, it's a vanity metric.
  2. Comparability — prefer rates and ratios (conversion, retention %) over raw counts: they compare across periods and segments.
  3. Hard to game — ask "how could this be inflated without creating value?". "Clicks" are easy to game; "activated users" are harder.
  4. A precise definition — one numerator, one denominator, a fixed window and filters, or two teams will count it differently.
Property Metric KPI
What it is any measurable number a metric with a goal
Tied to a target not necessarily always (target)
How many many a chosen few
Example time on page activation rate ≥ 40% by Q3

⚠️ Common mistake: calling everything a KPI. KPIs are the 3–5 metrics a team is actually judged on; the rest are diagnostic metrics.

How do you pick the right chart type: bar, line, scatter, or pie?

Short answer: The chart type is dictated not by taste but by the kind of relationship in the data: comparing categories — bars, change over time — a line, correlation between two variables — a scatter plot, share of a few parts in a whole — a pie. When torn between a pie and anything else, almost always pick bars.

In depth:

First state which relationship you are showing, and only then choose the visual:

Relationship in data Chart type Why
Comparing categories Bar the eye compares length precisely; axis must start at 0
Change over time Line a line emphasizes trend and continuity
Correlation of X and Y Scatter reveals the cloud, clusters and outliers
Parts of a whole Pie / stacked only for 2–4 parts, summing to 100%
Distribution Histogram / boxplot shows shape, spread, tails
  • Bars — the workhorse for categories; the Y axis starts at zero, otherwise the height difference lies.
  • Line — only for an ordered continuous axis (time, days); never connect unrelated categories with a line.
  • Scatter — for "are X and Y related"; add a trend line if needed.

⚠️ Common mistake: a pie chart with 8 slices — the shares become indistinguishable; replace it with horizontal bars sorted by value.

Ready to make it stick?

Explore a sample, choose a track, and build a review routine that fits your preparation.

Questions about this track

How should I prepare for a Data Analyst interview?

Study the concepts you'll be asked to explain, not just the ones you can code. RecallDeck's Data Analyst track gives you 265+ curated interview questions and resurfaces each one with an Anki-style SM-2 schedule based on your ratings — to practise recalling and explaining them. Combine reviews with coding and mock interviews.

What topics does the Data Analyst track cover?

The Data Analyst track is organised into the core areas Data Analyst interviews actually test, grouped by topic and by difficulty (Concept, Junior, Middle, Senior). You can preview the full outline and sample questions above before signing in.

Is spaced repetition effective for Data Analyst interview prep?

Yes. Active recall and spaced practice can help retain what you study. Grade yourself honestly and check your understanding with practical work. RecallDeck schedules each Data Analyst card using your ratings, so familiar cards return less often and difficult ones get more practice.

Can I try the Data Analyst track before paying?

Eligible new monthly or yearly subscribers can try the complete Data Analyst track and every feature for seven days. Checkout confirms trial eligibility and the amount due; returning subscribers may be charged immediately. Cancel online before the trial ends to avoid its first charge. Lifetime access has no trial.

Other interview tracks