Backend Engineer interview prep
A spaced-repetition deck of 625+ Backend Engineer interview questions — organised by topic and difficulty, and scheduled for review based on your ratings. Preview a few cards below, then choose access to study the whole track on an Anki-style SM-2 schedule.
Eligible new monthly or yearly subscribers get 7 days free. Checkout confirms eligibility and charges.
Before you start
Start with experience building an API and using a database. Shared foundations use Python, SQL, Django/FastAPI, and SQLAlchemy; backend modules add distributed-systems and reliability depth.
First pass: Python → Databases. Begin with the foundational questions in these modules, then follow the outline. Revisit advanced and senior questions as the job requires; all cards remain available from the start.
Shared topics can appear in several tracks and keep one review history. The outline labels shared foundations and role-specific modules; choose your depth for the job you are preparing for.
What's covered
Every topic in this track, grouped the way you'd study it.
Python
122 cardsShared foundations
Databases
104 cardsShared foundations
Backend
169 cardsShared foundations
CS Fundamentals
76 cardsShared foundations
DevOps & Infra
35 cardsShared foundations
System Design
29 cardsShared foundations
Distributed Systems
12 cardsRole-specific module
Kafka & Messaging
12 cardsRole-specific module
Caching Deep-Dive
10 cardsRole-specific module
Performance & Concurrency
11 cardsRole-specific module
Reliability & Incidents
10 cardsRole-specific module
Behavioral
35 cardsShared foundations
Sample questions
A few cards from the deck — reveal each answer, then choose access to study the full set on a schedule.
What's the difference between mutable and immutable types?
What's the difference between mutable and immutable types?
Short answer: Mutable objects can change in place (list, dict, set); immutable objects cannot (int, str, tuple). Rebinding a variable is different from mutating its object.
items = [1, 2]
alias = items
items.append(3)
assert alias is items and alias == [1, 2, 3]
text = "hello"
original = text
text += " world"
assert original == "hello" # the original string did not change
In depth:
id()identifies an object and stays constant during its lifetime. Treating it as a memory address is a CPython implementation detail.- A tuple cannot replace its elements, but a list stored inside it can still change. Immutability is not automatically deep immutability or thread safety.
- Dictionary keys and set elements must be hashable: their hash stays stable, and equal objects have equal hashes. A tuple containing a list is unhashable. A user-defined mutable instance can be hashable by identity, so hashability is not synonymous with immutability.
- Functions receive references to objects. Mutating an argument can affect the caller; rebinding a local parameter does not rebind the caller's variable.
Sources: Python hashability, object identity.
What does SELECT do, and in what order do the parts of a query execute?
What does SELECT do, and in what order do the parts of a query execute?
Short answer: SELECT retrieves rows from tables. The logical execution order does NOT match the written order: first FROM, then WHERE, GROUP BY, HAVING, SELECT, DISTINCT, ORDER BY, LIMIT.
In depth:
Written (syntactic) order:
SELECT DISTINCT col1, agg(col2)
FROM t
WHERE cond
GROUP BY col1
HAVING agg_cond
ORDER BY col1
LIMIT 10 OFFSET 20;
Logical execution order:
1. FROM / JOIN -- which tables, how to join
2. WHERE -- row filter BEFORE grouping
3. GROUP BY -- grouping
4. HAVING -- group filter AFTER aggregation
5. SELECT -- evaluate expressions, aliases
6. DISTINCT -- remove duplicates
7. ORDER BY -- sorting
8. LIMIT / OFFSET -- slice
⚠️ Gotcha: an alias from SELECT cannot be used in WHERE (since WHERE runs before SELECT), but it can be used in ORDER BY and often in GROUP BY (depends on the DBMS). Example of the error:
SELECT salary * 12 AS annual FROM emp WHERE annual > 100000; -- ERROR
SELECT salary * 12 AS annual FROM emp ORDER BY annual; -- OK
What is HTTP and how does the request-response cycle work?
What is HTTP and how does the request-response cycle work?
Short answer: HTTP (HyperText Transfer Protocol) is a text-based application-layer client-server protocol on top of TCP (in HTTP/3 — on top of QUIC/UDP). The client sends a request, the server returns a response; the connection holds no state between requests (stateless).
In depth:
A request consists of a start line (method + path + version), headers, and an optional body:
POST /api/v1/users HTTP/1.1
Host: api.example.com
Content-Type: application/json
Accept: application/json
Content-Length: 38
{"name": "Anna", "email": "a@ex.com"}
A response consists of a status line (version + code + reason phrase), headers, and a body:
HTTP/1.1 201 Created
Content-Type: application/json
Location: /api/v1/users/42
{"id": 42, "name": "Anna", "email": "a@ex.com"}
⚠️ Gotcha: "HTTP only works over TCP" is incorrect for HTTP/3, which uses QUIC over UDP. And the reason phrase ("OK", "Created") is purely informational — clients should rely on the numeric code, not the text.
What is Big O notation?
What is Big O notation?
Short answer: Big O describes how an algorithm's running time or memory usage grows as the input size n increases, dropping constants and lower-order terms. It's an upper bound on the growth rate.
In detail:
Big O answers the question "what happens when n becomes very large?". We're not interested in the exact number of operations, only the nature of the growth. That's why O(2n + 100) simplifies to O(n), and O(3n² + n) — to O(n²).
The main growth classes (from best to worst):
# O(1) — constant: doesn't depend on n
def first(arr):
return arr[0] if arr else None
# O(log n) — logarithmic: each step halves the problem (binary search)
def binary_search(arr, target):
lo, hi = 0, len(arr) - 1
while lo <= hi:
mid = (lo + hi) // 2
if arr[mid] == target:
return mid
elif arr[mid] < target:
lo = mid + 1
else:
hi = mid - 1
return -1
# O(n) — linear: a single pass
def total(arr):
s = 0
for x in arr: # n iterations
s += x
return s
# O(n log n) — efficient sorts (merge, quick, Timsort)
def sort_it(arr):
return sorted(arr)
# O(n^2) — quadratic: a nested loop
def has_dup_naive(arr):
for i in range(len(arr)):
for j in range(i + 1, len(arr)):
if arr[i] == arr[j]:
return True
return False
# O(2^n) — exponential: naive Fibonacci, enumerating subsets
def fib_naive(n):
if n < 2:
return n
return fib_naive(n - 1) + fib_naive(n - 2)
A rough guide to growth at n = 1,000,000: O(1) — 1 operation, O(log n) — ~20, O(n) — a million, O(n log n) — ~20 million, O(n²) — a trillion (already too much), O(2^n) — infeasible even at n = 50.
⚠️ Gotcha: Big O is about asymptotics (behavior at large n), not about actual time. An O(n) algorithm can be slower than an O(n²) one on small data because of large constants. People also confuse: Big O (upper bound), Ω (lower bound), Θ (tight bound); in interviews "O" usually means Θ.
What is Docker and what problem does it solve?
What is Docker and what problem does it solve?
Short answer: Docker packages an application and its user-space dependencies in an image, then runs an isolated process from that image. This reduces differences between development, CI and production.
For a Go service, the image can contain the compiled binary, CA certificates and any needed shared libraries. The host still supplies a compatible kernel and CPU architecture; configuration, credentials, resource limits and external services are supplied separately.
docker build -t myapp:1.0 .
docker run -d -p 127.0.0.1:8000:8000 --name myapp myapp:1.0
docker logs -f myapp
The application must listen on the container interface, such as :8000. This example publishes only to the host loopback address for local use.
Limit: the same image does not guarantee identical behavior on every machine. Linux containers on macOS/Windows normally run inside a Linux VM; CPU architecture, kernel features, networking and external configuration still matter. Build the required platform variants and test the actual deployment environment.
Sources: Reference 1, Reference 2.
Event sourcing and CQRS: when are they justified and what is the price?
Event sourcing and CQRS: when are they justified and what is the price?
Short answer: Event sourcing: the event log is authoritative; current state is derived as state = fold(events). CQRS: the write model and the read models are separated. You get a full audit trail, temporal queries ('what did the order look like yesterday'), and rebuildable projections. The price is steep: event versioning, snapshots, potential projection lag, expensive tooling. It is not a default architecture.
In depth:
events: OrderCreated → ItemAdded → ItemAdded → OrderPaid
state = fold(events) — always derivable from scratch
projections: 'orders per day', 'top items' — separate read
models, rebuildable from the log at any time
- When justified — money movements, ledgers, audit-heavy and compliance-bound domains; 'why did the balance end up like this' is a business question, not a log-digging exercise.
- Price #1: versioning — an event lives forever; when the schema changes you must read every old version (upcasters) or migrate the whole log.
- Price #2: reads — if read models update asynchronously, a projection can lag after a command; the UI, tests, and support must handle that consistency model.
- CQRS without ES — legitimate and far cheaper: a regular write DB + denormalized read projections.
⚠️ Common mistake: proposing event sourcing for a CRUD app 'for future growth'. If audit is not a business requirement, you pay the full ES price and gain nothing a table plus a change log would not give you.
The event log is authoritative, but implementations can store snapshots and materialized state. CQRS separates read and write models; it need not use separate databases, and asynchronous projection lag is a design choice rather than a requirement of every CQRS implementation.
Ready to make it stick?
Explore a sample, choose a track, and build a review routine that fits your preparation.
Questions about this track
How should I prepare for a Backend Engineer interview?
Study the concepts you'll be asked to explain, not just the ones you can code. RecallDeck's Backend Engineer track gives you 625+ curated interview questions and resurfaces each one with an Anki-style SM-2 schedule based on your ratings — to practise recalling and explaining them. Combine reviews with coding and mock interviews.
What topics does the Backend Engineer track cover?
The Backend Engineer track is organised into the core areas Backend Engineer interviews actually test, grouped by topic and by difficulty (Concept, Junior, Middle, Senior). You can preview the full outline and sample questions above before signing in.
Is spaced repetition effective for Backend Engineer interview prep?
Yes. Active recall and spaced practice can help retain what you study. Grade yourself honestly and check your understanding with practical work. RecallDeck schedules each Backend Engineer card using your ratings, so familiar cards return less often and difficult ones get more practice.
Can I try the Backend Engineer track before paying?
Eligible new monthly or yearly subscribers can try the complete Backend Engineer track and every feature for seven days. Checkout confirms trial eligibility and the amount due; returning subscribers may be charged immediately. Cancel online before the trial ends to avoid its first charge. Lifetime access has no trial.