This page collects 27 Python interview questions with model answers, grouped from fundamentals to architecture, plus red flags and a practical exercise. They come from how we evaluate senior Python engineers at Ryz: recruiters source people with real production Python behind them, candidates go through structured NTRVSTA AI interviews on internals, design and debugging, and recruiters review every candidate before and after. AI scores are advisory. People make the hiring decisions.
Pick six to eight questions that match the level and the kind of Python the role involves. Fundamentals tell you whether someone understands the language they write. The senior section tells you whether they can own a service in production.
Listen for reasoning, not recall. A candidate who says "I'd measure first with py-spy, then decide" is more useful than one who recites the MRO algorithm. Pair the conversation with a short practical exercise, since talking about Python and writing it are different skills.
Default argument values are evaluated once, when the def statement runs, so a list default is shared across calls. Closures look up free variables when the inner function runs, not when it is created, so lambdas built in a loop all see the loop variable's final value.
def add(item, bucket=None):
if bucket is None:
bucket = []
bucket.append(item)
return bucket
handlers = [lambda x, i=i: x * i for i in range(3)] # bind i now
What a strong answer shows: They explain when each binding happens, not just the fix, and mention linters (Ruff's B006/B023 rules) that catch both in review.
is and == differ, and when is is the right choice?== calls __eq__ and compares values. is compares identity. Use is for singletons such as None, sentinel objects and enum members.
What a strong answer shows: They know interning is a CPython implementation detail and that a custom __eq__ can make x == None lie.
__eq__ on a class and suddenly instances cannot go into a set. Why?Defining __eq__ without __hash__ sets __hash__ to None, because the default identity hash would break the rule that equal objects hash equally. Fix it by defining __hash__ over the same immutable fields, or use @dataclass(frozen=True), which generates both.
What a strong answer shows: They connect hashing to the equality contract and warn that hashing mutable fields breaks dict lookups once a field changes.
A generator function returns an iterator that produces values lazily and keeps its frame suspended between yields. Use one when the data is large or unbounded, when you want a pipeline of transformations without intermediate lists, or when the consumer may stop early.
def read_events(path):
with open(path) as f:
for line in f:
if line.strip():
yield json.loads(line)
errors = (e for e in read_events("app.log") if e["level"] == "error")
What a strong answer shows: They mention that a generator can only be consumed once and that the file stays open until the generator finishes or is closed.
It guarantees setup and teardown around a block, with teardown running even if the block raises. Write a class with __enter__ and __exit__, or a generator decorated with contextlib.contextmanager with the cleanup in a finally. ExitStack handles a dynamic number of resources.
What a strong answer shows: They put cleanup in finally, know __exit__ can suppress exceptions by returning true, and say why that is rarely wanted.
NamedTuple, a TypedDict and a Pydantic model?Dataclasses are plain internal value objects with optional frozen and slots. NamedTuple suits small immutable records that must behave like tuples. TypedDict types dictionaries you do not control, such as JSON from an API, with no runtime cost. Pydantic validates and coerces untrusted input at boundaries, and is the right choice for request bodies and settings.
What a strong answer shows: They validate at the edges with Pydantic and keep the core on cheap dataclasses, instead of using one tool everywhere.
Good answers name a few concrete features: structural pattern matching (3.10), ExceptionGroup and asyncio.TaskGroup (3.11), the type statement and inline generic syntax def first[T](xs: list[T]) -> T (3.12), the experimental free-threaded build and new interactive REPL (3.13), and deferred evaluation of annotations plus template strings (3.14).
What a strong answer shows: They tie features to code they wrote and know which version their production fleet runs, not just the release notes.
Importing a module executes it top to bottom. If a imports b and b imports a name from a that has not been defined yet, you get an ImportError on a partially initialized module. Quick fixes move the import into a function or behind if TYPE_CHECKING:. The real fix is usually structural: extract the shared piece into a third module so dependencies point one way.
What a strong answer shows: They treat the cycle as a design signal and mention tooling such as import-linter to enforce layering.
Confirm with django-debug-toolbar or query logging in a test using assertNumQueries. The usual cause is an N+1 in a template loop. Use select_related for foreign keys and one-to-ones (a join), prefetch_related for reverse and many-to-many relations (a second query), Prefetch objects to filter the prefetched set, and annotate with Count instead of calling .count() per row.
What a strong answer shows: They add a query-count test so the regression cannot return, and check the generated SQL and indexes rather than assuming the ORM did the right thing.
functools.wraps?A decorator is a callable that takes a function and returns a replacement. Without functools.wraps, the wrapper hides the original's __name__, docstring, __wrapped__ and signature, which breaks introspection used by FastAPI, pytest fixtures and debuggers.
import functools, time
def timed(fn):
@functools.wraps(fn)
def wrapper(*args, **kwargs):
start = time.perf_counter()
try:
return fn(*args, **kwargs)
finally:
log.info("%s took %.3fs", fn.__qualname__, time.perf_counter() - start)
return wrapper
What a strong answer shows: They note that this wrapper breaks for async def functions and would branch on inspect.iscoroutinefunction, and they type it with ParamSpec.
Run mypy or Pyright in CI with strictness ratcheted per package, type the boundaries first (public functions, API schemas, database models), and use Protocol for structural interfaces instead of deep inheritance.
What a strong answer shows: A migration plan that does not stop feature work, and an understanding that hints are not enforced at runtime unless something like Pydantic does it.
Use a session-scoped fixture to start the database (often Testcontainers), a function-scoped fixture that wraps each test in a transaction and rolls it back, and factories for test data. Mock the external API at the HTTP layer with respx or responses, and keep a small set of contract tests that hit a sandbox.
What a strong answer shows: They balance speed and fidelity, and avoid mocking their own database layer so tests still catch SQL bugs.
__getattr__ and __getattribute__, and where do descriptors fit?__getattribute__ runs on every attribute access. __getattr__ runs only when normal lookup fails. Descriptors (objects with __get__, __set__ or __delete__) are how property, methods, classmethod and ORM fields work.
What a strong answer shows: They can explain how a Django or SQLAlchemy model field turns attribute access into a query, and they avoid overriding __getattribute__ because of infinite recursion and speed.
The likely cause is blocking work on the event loop: a synchronous HTTP client, a sync database driver, heavy JSON or Pydantic work, or CPU-bound code inside an async def route. Every request shares one loop per worker, so one blocking call stalls all of them. Turn on asyncio debug mode to log slow callbacks, profile with py-spy, then switch to async clients, declare blocking routes as plain def so they run in the threadpool, or offload with asyncio.to_thread.
What a strong answer shows: They know FastAPI runs sync routes in a threadpool and async routes on the loop, and they check connection pool sizes before blaming Python.
The GIL lets only one thread execute Python bytecode at a time in a process, so threads help with I/O-bound work but not CPU-bound pure Python. It does not make your code thread-safe: compound operations like counter += 1 can still interleave. The free-threaded build (experimental in 3.13 and officially supported in 3.14, still opt-in) removes the GIL, at some single-threaded cost and with C extensions needing to declare support.
What a strong answer shows: They would benchmark their real workload and check that key native dependencies ship free-threaded wheels before switching, and still reach for processes or vectorized libraries for CPU work today.
Use asyncio.TaskGroup so a failure cancels siblings and errors surface as an ExceptionGroup, a Semaphore to cap concurrency, and asyncio.timeout for deadlines.
async def fetch_all(client, urls, limit=20):
sem = asyncio.Semaphore(limit)
async def fetch(url):
async with sem, asyncio.timeout(10):
r = await client.get(url)
r.raise_for_status()
return r.json()
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(fetch(u)) for u in urls]
return [t.result() for t in tasks]
What a strong answer shows: They explain why bare create_task without holding a reference is risky, and they never swallow CancelledError.
Asyncio suits many concurrent network connections in one process. Threads suit blocking libraries you cannot replace and moderate I/O concurrency. Processes, via ProcessPoolExecutor or a job queue, suit CPU-bound pure Python on the standard build. Libraries like NumPy and Polars release the GIL in native code, so threads can scale there.
What a strong answer shows: They factor in pickling overhead and memory per process, and mention subinterpreters or the free-threaded build as options they would evaluate, not default to.
Profile first to find whether time goes to I/O, row-wise apply calls, or memory pressure and swapping. Common wins: vectorize instead of apply, read only needed columns from Parquet, use categorical and smaller numeric dtypes, push filters into the source query, or move to Polars' lazy API so the query optimizer can stream and prune.
What a strong answer shows: They measure memory as well as CPU and question whether the job needs the whole dataset in memory at all.
Watch RSS over time to confirm growth, then use tracemalloc snapshots and compare top allocations, or memray for native allocations. Usual suspects: unbounded caches like a module-level dict or lru_cache on methods (which holds self), reference cycles holding large objects, and native extensions.
What a strong answer shows: They distinguish Python-level leaks from allocator fragmentation and say how they would reproduce it outside production.
True exactly-once delivery is not available. Aim for at-least-once delivery with idempotent processing. Store an idempotency key per charge in Postgres with a unique constraint, record state transitions in the same transaction as the business change, pass the key to the payment provider's idempotency API, and enable late acknowledgement so crashed tasks are redelivered.
What a strong answer shows: They name the failure windows (crash after charge, before commit) and reconcile against the provider instead of trusting retries.
Separate installable packages with explicit dependencies, a workspace tool such as uv workspaces or Pants, enforced import boundaries, one lockfile strategy, and CI that only tests what changed.
What a strong answer shows: They focus on ownership and dependency direction, not folder names, and plan for test time as the repo grows.
Map the domain and data ownership first. Extract along a boundary with few writes crossing it, put an internal interface in place inside the monolith before moving code, and use the strangler pattern with traffic routing.
What a strong answer shows: They question whether splitting is needed at all, and know that a modular monolith often solves the team problem with less operational cost.
Treat the model as a slow, unreliable dependency: async client, timeouts, retries with backoff for rate limits, streaming responses to the user, and a queue for long jobs. Keep prompts versioned in code, validate structured outputs with Pydantic, log inputs and outputs with PII handling, and run an evaluation set in CI before changing prompts or models.
What a strong answer shows: They treat evaluation and cost tracking as part of the design and plan for the model returning malformed output.
Pin with a lockfile, upgrade one canary service first, run the full test suite plus warnings-as-errors for deprecations, deploy behind a gradual rollout, and watch error rates and latency.
What a strong answer shows: A repeatable process with rollback, and awareness of transitive dependencies that lag behind.
Structured JSON logs with request IDs, OpenTelemetry traces across HTTP, database and queue calls, RED metrics per endpoint, and alerts on symptoms users feel.
What a strong answer shows: They connect telemetry to questions they need answered during an incident and watch label cardinality.
Parameterized queries only, strict input validation, no pickle or yaml.load on untrusted data, secrets from a manager rather than env files in images, dependency scanning with pip-audit, least-privilege database roles, and SSRF protection on any endpoint that fetches URLs.
What a strong answer shows: They name Python-specific risks like deserialization and know how their framework handles CSRF and auth.
Only after profiling shows a CPU-bound hot spot that vectorization, better algorithms or an existing native library cannot fix. PyO3 with maturin is the common route today. Keep the boundary coarse so you cross it rarely, and release the GIL inside long native work.
What a strong answer shows: They weigh build and hiring cost against the gain and try simpler options first.
requests or time.sleep inside async def code and cannot explain why it hurts.except: or except Exception: pass as error handling.Give a 3-hour take-home: a small FastAPI service that ingests webhook events from a fake payment provider, verifies an HMAC signature, stores events in Postgres or SQLite, deduplicates by event ID, and exposes a paginated endpoint to list them. Provide a script that sends duplicate, out-of-order and malformed events. Ask for tests and a short README describing trade-offs and what they would do with more time. Then spend 30 minutes extending it live.
If you would rather skip building this loop yourself, hire senior Python developers through Ryz. They are the top 1% of the candidates we interview, they work on your team and in your repos, and they work within ±1h of US time zones. See how we assess them on our vetting process page, or start from our Python developer job description if you are hiring directly.
Six to eight in a 60-minute conversation is enough if you follow up on each one. Two or three well-explored questions about async behavior or data modeling reveal more than fifteen quick definitions.
Use them sparingly. Algorithm puzzles show problem solving but say little about packaging, testing, profiling or framework knowledge. A realistic service exercise or a debugging session on real code predicts day-to-day performance better.
Mid-level engineers usually know the right tool. Senior engineers explain when it stops being the right tool, what they would measure, and how they would roll out a change safely across a team and a production fleet.
Questions we didn't answer? Email info@ryzlabs.com.