For AI, machine learning and data science interviews

Code is written by AI now.
Interviews now test your judgment.

Land your dream data job. Practice with an AI interviewer built for data scientists, ML and AI engineers, and learn the part AI can’t do for you.

Why it matters

Anyone can generate the code.
The interview is about defending it.

Track What the AI writes in seconds What the interviewer actually asks
01Concepts
X_train, X_test = train_test_split(
    df, test_size=0.2)
The rows are three years of customer history. Is a random split the right call?
02System design
scores = model.predict(user_item_pairs)
top_k = scores.argsort()[-10:]
Recommendations must load in 50 ms for 10 million users. What do you compute ahead of time?
03Troubleshooting
df["income"] = df["income"].fillna(
    df["income"].median())
Most missing incomes come from one region. What did the median just do to your model?
04Product sense
dau = events.groupby("date")["user_id"].nunique()
dau.pct_change().tail(7)
Daily active users fell 10% last week. Is it a real problem, and how would you find out?
Inside a Pondly session

It follows your reasoning. Then it pushes back.

Troubleshooting · A RAG system starts giving wrong answersTough
Interviewer
Your RAG system suddenly starts giving incorrect answers. What’s the first thing you investigate, and how would you prove that’s actually the root cause?
Me
I’d split retrieval from generation first. Pull a sample of the bad answers and check whether the right documents were in the top results. If they weren’t, it’s retrieval, not the model.
Interviewer
Separating retrieval from generation is a strong first move: it halves where the fault can be. Say the right chunks are coming back. What changed recently, and how would you show it’s the cause rather than a coincidence?
Topics

Every topic they’ll push you on.

  • Class imbalance
  • Bias-variance trade-off
  • Confidence intervals
  • A/B testing
  • p-values
  • Simpson's paradox
  • Multiple testing correction
  • Bootstrapping
  • Calibration
  • Precision and recall
  • ROC AUC
  • Cross-validation
  • Hypothesis testing
  • Correlation vs causation
  • Uplift modelling
  • Difference-in-differences
  • RAG
  • Vector databases
  • Hallucination
  • Evaluating LLM output
  • Prompt injection and guardrails
  • Fine-tuning vs RAG
  • Chunking strategies
  • Reranking
  • Embeddings
  • Transformers
  • Agents
  • Temperature and sampling
  • Tokens and the context window
  • Structured output and function calling
  • Vector search vs keyword search vs hybrid
  • Debug a broken RAG system
  • Data drift vs concept drift
  • Data leakage
  • Train/serve skew
  • Feature importance
  • Gradient boosting
  • Missing values
  • Batch vs online prediction
  • Deployment strategies
  • Retraining triggers
  • Choosing a classification threshold
  • Regularisation
  • Ridge vs Lasso vs elastic net
  • Multicollinearity
  • Feature selection
  • Overfitting and underfitting
  • Baseline models
How it works

From question to debrief in three steps.

1Pick a question

Choose one, or let Pondly pick. Set the difficulty; it stays the same for the whole question.

2Talk it through

The interviewer asks, listens and follows up. Stuck? Hints build from a few ideas to the full answer.

3Read the debrief

Four ratings with evidence from your own words, one thing to keep, one habit to sharpen.