Prepare a decision you can substantiate
These 15 questions and 30 follow-ups are original preparation exercises, grounded in official facts checked on 2026-10-05. They are not questions Databricks has confirmed it asks, and the “what it may test” text is editorial inference, not a scoring rubric. Read the role-and-interview facts article first to resolve the fact IDs.
The general Tokyo FDE postings emphasize production implementation, Spark runtime knowledge and multiple clouds. The Tokyo AI FDE posting emphasizes GenAI applications, evaluation and production ML on one cloud. Each exercise states its applicable role; where both roles appear, it preserves which fact belongs to which. Production reasoning also matters at other employers. The shared discipline is not evidence of a shared hiring process.
Use your own experience. If you have not performed a task, say so and use a labeled offline design exercise. Metrics below are types of evidence to supply when relevant; they are neither fabricated achievements nor Databricks acceptance thresholds. No customer credentials, paid calls or live workspace are needed.
- 1Real project or labeled exercise
- 2Data and responsibility boundaries
- 3Failure hypothesis
- 4Controlled comparison
- 5Release or recovery decision
- 1Missing evidence
- 2Small synthetic case
- 3Limits recorded
- 4Next verification plan
This original preparation flow draws its work areas from DB005, DB011–DB012, DB017 and DB021–DB023. It is not an internal architecture or interview sequence. Follow it by keeping one case stable while changing a hypothesis; the final answer should show why a decision was made, not only the tools used.
DB-TQ01: How would you trace one customer request from source data through a model to its user interface?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB002, DB005, DB012, DB017 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002 and Tokyo FDE CSQ427R277 / ATS 8741922002. DBS01 official source, DBS02 official source.
What it may test — inference: Whether you can explain the complete production path, including boundaries owned by other teams, rather than presenting the model as the whole application.
Evidence, numbers and decisions to include: Choose an actual project and draw data arrival, transformation, model invocation, response and UI. Identify the revision, identity and request ID passed between components. State your own implementation and decisions separately from team work. Use observed input volume, freshness lag, error rate and end-to-end p95 latency with units and a measurement window. Show a rejected design and the constraint that ruled it out.
Deep dive 1: The model responds correctly, but the user sees stale data. What would you inspect?
Points to explain: Compare source revision, pipeline completion, cache/index age and UI state for the same request. Explain how you distinguish stale inputs from stale presentation; name the owner and recovery action for each boundary. Do not call model retraining the first remedy.
Deep dive 2: A failed request is retried. How do you prevent an inconsistent user result?
Points to explain: Explain stable request identifiers, which operations can be repeated, duplicate detection and the visible pending/failed state. If a write is involved, separate authorization and commit from answer generation. Describe the test that proves the response belongs to the intended data revision.
Avoid a shallow answer: “I connected the API and the UI.” That leaves data revision, responsibility and recovery unexplained.
DB-TQ02: A production Spark job becomes much slower. Which hypotheses would you test first?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB007, DB011, DB012 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002. DBS01 official source.
What it may test — inference: Whether you connect distributed-runtime behavior to evidence, instead of changing compute size until the job finishes.
Evidence, numbers and decisions to include: Bring a sanitized execution plan and stage/task observations from a real slowdown. Compare input size and distribution, task-duration spread, shuffle, spill, I/O and executor failures against the same workload’s baseline. Explain the order of hypotheses and why each observation supports or rejects one. Report runtime, retry behavior and resource cost before and after one controlled change.
Deep dive 1: How would you distinguish skew from broad resource saturation?
Points to explain: Contrast a small number of long tasks with slowdown across many tasks. Inspect the data distribution and partitioning associated with the affected stage. State what further evidence would falsify your hypothesis; explain why adding workers might leave the long tail unchanged.
Deep dive 2: The job is faster after tuning. What could still make the change unsuitable?
Points to explain: Check output correctness, peak load, cost, failure rate and sensitivity to different input distributions. Record the configuration and workload used, and describe a regression check. Avoid treating a single warm-cache run as a repeatable production improvement.
Avoid a shallow answer: “I increased the cluster.” The answer needs a cause, a controlled comparison and the remaining trade-off.
DB-TQ03: How would you preserve a pipeline’s correctness when data is late, duplicated or reprocessed?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB011, DB012, DB017 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002 and Tokyo FDE CSQ427R277 / ATS 8741922002. DBS01 official source, DBS02 official source.
What it may test — inference: Whether your production design accounts for data semantics and recovery, beyond achieving acceptable job runtime.
Evidence, numbers and decisions to include: Describe an actual data contract: entity key, event/arrival time, expected ordering and correction policy. Identify where duplicate or late records were detected and how a replay changed downstream model inputs and UI results. Use actual record counts, reconciliation differences and data-lag measurements. Distinguish a delivery guarantee you verified from one merely assumed.
Deep dive 1: A retry happens after some output was written. What must the recovery prove?
Points to explain: Show the commit boundary and the evidence that identifies completed work. Explain whether a retry overwrites, appends or reconciles records and how downstream readers observe a consistent version. Test partial progress, not only a clean restart.
Deep dive 2: A business correction arrives after a report was used. What would you change?
Points to explain: Define who authorizes the correction, what period is restated and how consumers learn the revision. Explain the difference between preserving an audit trail and silently replacing a past answer. Give a reconciliation procedure and its owner.
Avoid a shallow answer: “The platform handles retries.” That does not establish the business meaning of duplicates or corrections.
DB-TQ04: When a customer asks for another cloud, what can you reuse and what must you verify again?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB005, DB010, DB024 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002; comparison with Tokyo AI FDE FEQ227R196 / ATS 8569392002. DBS01 official source, DBS03 official source.
What it may test — inference: Whether you distinguish application logic from cloud-specific identity, network, storage and operations; this is a practice inference from the two-cloud FDE requirement.
Evidence, numbers and decisions to include: Choose two clouds you actually know and state which one is your specialist area. Map the application contract, identity flow, data boundary, connectivity, deployment and monitoring separately. Give a small verification plan for the unfamiliar boundary and a rollback condition. Use measured migration effort or known constraints if available; label unexecuted comparisons. The AI FDE posting requires production ML on one listed cloud, so do not claim its requirement is also two clouds.
Deep dive 1: The customer’s identity and network controls conflict with your preferred architecture. What changes?
Points to explain: Identify the enforceable boundary and the security owner. Compare a smaller integration with a full migration using data movement, operational responsibility and failure recovery. Explain which validation must precede production access.
Deep dive 2: How do you avoid claiming portability you have not demonstrated?
Points to explain: List what is unchanged and what has been tested on each cloud, with versions and workload. Separate design equivalence from measured behavior. State the missing evidence and who can review the cloud-specific part.
Avoid a shallow answer: “The cloud services are equivalent.” Similar names do not prove equivalent controls or operations.
DB-TQ05: How would you release a change that affects code, data processing and model behavior together?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB012, DB017 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002 and Tokyo FDE CSQ427R277 / ATS 8741922002. DBS01 official source, DBS02 official source.
What it may test — inference: Whether CI/CD and MLOps are a reproducible release path with an accountable rollback, rather than names of tools.
Evidence, numbers and decisions to include: Use a release you actually owned. Identify code, dependencies, data/schema, model or prompt versions, configuration and the environment promoted. Show test gates and the evidence retained at each gate. Report the actual deployment window, failure criteria and recovery time if measured. Explain who can approve a release and why a previous version remains usable.
Deep dive 1: Code rollback succeeds, but the data schema has changed. What now?
Points to explain: Describe the irreversible or separately reversible steps. Explain how you protect readers during the transition, which versions remain compatible and what recovery was tested. Do not assume rolling back a binary restores data.
Deep dive 2: A model change passes unit tests but shifts output quality. What would block release?
Points to explain: Use a versioned evaluation set with relevant failure groups, not only code tests. Explain the agreed acceptance conditions, monitored canary population and stop/rollback owner. Keep hypothetical thresholds separate from actual project thresholds.
Avoid a shallow answer: “We used CI/CD.” Name the artifacts, gates and failure path that made the release reproducible.
DB-TQ06: How would you design a secure data-and-AI application when users have different access rights?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB005, DB017 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002 and Tokyo FDE CSQ427R277 / ATS 8741922002. DBS01 official source, DBS02 official source.
What it may test — inference: Whether you preserve a user’s authority across data retrieval, model calls and UI results; the specific design is editorial inference, not a published interview task.
Evidence, numbers and decisions to include: Sketch a real or explicitly hypothetical request with a principal, permitted data, service identity and any write operation. Explain where authorization is enforced and what the model cannot decide. Show tests for an unauthorized user, changed permissions, empty evidence and unsafe output. State what is logged, who can view it and which sensitive content you intentionally exclude.
Deep dive 1: The application returns only a summary. Could it still expose restricted data?
Points to explain: Trace all material supplied to the model, including retrieval results and logs. Explain how access is constrained before generation and how the response is checked. A short summary does not remove the data boundary.
Deep dive 2: An agent requests an action with the application’s service identity. Who authorizes it?
Points to explain: Separate the requesting user from the executing service. Validate the operation and resource against deterministic policy; describe confirmation for consequential changes when appropriate. Show denial and audit evidence without exposing credentials.
Avoid a shallow answer: “The model is instructed not to disclose.” A prompt alone does not prove authorization.
DB-TQ07: What evidence would make you comfortable releasing a GenAI application to its intended users?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB021, DB023 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether evaluation connects intended use, consequential failure and a production decision, instead of treating an average score as sufficient.
Evidence, numbers and decisions to include: Define the application’s users, decisions and failure consequences. Bring a versioned evaluation set with expected outcomes and representative failure categories. State who selected the cases and approved acceptance conditions. Use actual sample counts, per-category failure rates, latency and cost where measured. Explain limitations, rollout population, monitoring and the person able to stop the release.
Deep dive 1: The average score improves while a critical failure group gets worse. Would you ship?
Points to explain: Show the group’s consequence and coverage rather than averaging it away. Compare actual acceptance conditions, uncertainty from sample size and possible scope restrictions. Explain whether you hold, restrict or release, and what evidence changes that decision.
Deep dive 2: Offline results are good, but users abandon the application. What would you inspect?
Points to explain: Separate task success from evaluation score. Compare the live input mix, UI friction, latency, evidence usefulness and user workflow. Explain how reviewed real failures enter the next evaluation revision without treating model output as ground truth.
Avoid a shallow answer: “Accuracy exceeded our target.” Define the target, failures, sample and production decision.
DB-TQ08: How would you identify whether a wrong RAG answer came from retrieval or generation?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB022, DB023 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether you can isolate a failure in an experience area explicitly listed by the AI FDE posting.
Evidence, numbers and decisions to include: Bring an actual failed case or a labeled synthetic case with question, permitted source revision, retrieved passages and output. Explain the expected evidence and whether it was available, selected and used correctly. Compare retrieval coverage and grounded-answer quality separately. Record actual corpus size, freshness boundary, sample counts and the effect of one change; do not claim a larger model solved a retrieval problem without observing it.
Deep dive 1: The right passage is retrieved, but the answer is still wrong. What next?
Points to explain: Inspect context assembly, conflicting revisions, question interpretation and citation alignment. Change one part and rerun the same cases. Include a case where evidence is insufficient and the correct behavior is to hold.
Deep dive 2: Improved retrieval raises quality but increases latency and cost. How do you decide?
Points to explain: Measure the added material and time per task, plus the failure groups improved. Compare narrower retrieval, reranking or a smaller scope as candidate experiments. State the actual acceptance constraint and evidence required before changing the design.
Avoid a shallow answer: “We added embeddings and RAG.” Describe where a failure originated and which comparison justified a fix.
DB-TQ09: When would you choose retrieval, fine-tuning or a simpler deterministic solution?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB022, DB023 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether listed techniques are alternatives chosen for a failure mode, rather than a checklist used to make a design look sophisticated.
Evidence, numbers and decisions to include: Start with a real task and baseline. Distinguish missing current facts, unstable output behavior and a task whose answer can be computed directly. Compare data freshness, training/evaluation data rights, change frequency, operating effort and quality. Show measured results if you ran alternatives; otherwise present an unexecuted comparison plan and the evidence you lack.
Deep dive 1: Fine-tuning improves the sample, but the customer’s facts change daily. What does that leave unresolved?
Points to explain: Explain which behavior was changed and where fresh facts still come from. Test stale or missing evidence separately from desired output style. State the version/update responsibility and what invalidates the claimed improvement.
Deep dive 2: A simpler SQL or rules solution performs adequately. Why keep an LLM?
Points to explain: Identify the user need that remains unsolved and its measurable consequence. If none remains, justify the simpler option. Compare operational cost and failure recovery, rather than arguing from the customer’s request for AI.
Avoid a shallow answer: “Fine-tuning is more accurate” or “RAG is always better.” State the task, failure and comparison.
DB-TQ10: What would justify a multi-agent design over one agent with a small set of tools?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB021, DB022, DB023 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether added coordination solves a demonstrated problem and remains evaluable in production.
Evidence, numbers and decisions to include: Describe a task boundary, evidence available to each component, permitted actions and the final decision owner. Compare a single-agent baseline with the proposed split using actual success/error rates, tool-call counts, latency and recovery effort if measured. Specify message contracts, stopping conditions and what a trace must show. Label the design hypothetical if you have not implemented it.
Deep dive 1: Two agents disagree or repeat the same action. What prevents an uncontrolled loop?
Points to explain: Explain bounded execution, duplicate-action protection and the authority to resolve disagreement. Show the observable state and a failure test. More agents do not justify letting a model decide permissions or irreversible commits.
Deep dive 2: A component performs well alone but the full workflow degrades. How would you evaluate it?
Points to explain: Keep component and end-to-end cases separate. Inspect handoff information loss, conflicting evidence, accumulated latency and error propagation. State which boundary you would remove or simplify if it adds no verified value.
Avoid a shallow answer: “Each agent has a specialist role.” Give evidence that the split improves the complete task.
DB-TQ11: How would you verify a Text2SQL answer when the business term and access boundary are both ambiguous?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB022, DB023, DB005 — Tokyo AI FDE FEQ227R196 / ATS 8569392002 for Text2SQL; Tokyo Sr. FDE CSQ327R52 / ATS 8568173002 for secure architecture. DBS03 official source, DBS01 official source.
What it may test — inference: Whether you treat SQL execution, business meaning and authorization as separate checks. Text2SQL is an AI FDE experience example; the secure-architecture fact belongs to general FDE.
Evidence, numbers and decisions to include: Choose a business question and define its term, time boundary, authorized views and expected result. Use a fixed data snapshot and cases for duplicate names, nulls, joins, forbidden columns and insufficient evidence. Show the query, rows/revision, answer and citation together. Record execution bounds and actual failure categories; do not equate a successful query with a correct answer.
Deep dive 1: The SQL is valid but the customer disputes the metric. How do you investigate?
Points to explain: Resolve the denominator, filters, time zone and business definition with its owner. Compare the expected result on the same snapshot. State whether the error is semantic, data-related or query generation, and version the corrected definition.
Deep dive 2: A user asks for a field outside their permission. What should the system do?
Points to explain: Enforce policy before query results reach the model. Explain explicit denial or a permitted narrower result, with no hidden inference from restricted rows. Demonstrate the denied case and which audit fields are safe to retain.
Avoid a shallow answer: “The query ran successfully.” Execution does not prove meaning, evidence or access compliance.
DB-TQ12: How would you reduce GenAI latency or cost without hiding a quality regression?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB021, DB023 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether optimization is constrained by user outcomes and relevant failure groups, rather than token price or a fastest isolated call.
Evidence, numbers and decisions to include: Define the task boundary and measure model time, retrieval/tool time, retries, context size and user-visible end-to-end latency separately. Use real task counts, p50/p95, cost per completed task and evaluation results with a measurement window. Compare one candidate change at a time, including the operating effort introduced. Do not invent a savings percentage for a system you have not measured.
Deep dive 1: A cheaper model needs more retries and human correction. Is it cheaper?
Points to explain: Count the complete task, including retries, failed attempts and review effort using an explicit accounting boundary. Separate measured costs from estimated human time. Explain whether the added error is tolerable for the task.
Deep dive 2: Caching improves latency, but the underlying data or permissions change. What must be checked?
Points to explain: State the cache key’s data revision and authority assumptions, invalidation rule and scope. Test stale evidence and changed permissions. Explain when avoiding the cache is preferable to returning an answer outside its valid boundary.
Avoid a shallow answer: “The model costs less per token.” The user pays for a completed, acceptable task.
DB-TQ13: What did you have to control to turn a cloud ML experiment into a production service?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB024, DB025, DB027 — Tokyo Senior AI Engineer–FDE FEQ227R196 / ATS 8569392002. DBS03 official source.
What it may test — inference: Whether your one-cloud production ML experience is operationally concrete; the preferred Databricks/Spark background is not silently made mandatory.
Evidence, numbers and decisions to include: Choose a cloud you actually used. Explain training/input lineage, artifact versions, serving capacity, identity, monitoring and the customer’s operating boundary. Use observed workload, deployment frequency, drift/failure observations and recovery measures. Identify the part you owned. If your experience is on another platform, map the concepts honestly and state which Databricks-specific behavior you still need to learn.
Deep dive 1: Training quality remains stable but live behavior changes. What would you compare?
Points to explain: Compare input distribution, feature/data revision, dependency and serving configuration against the deployed artifact. State what can be reproduced offline and which production observation is missing. Separate retraining from fixing a data or serving defect.
Deep dive 2: Would you reimplement a working customer ML service on a new platform?
Points to explain: Compare the required outcome, migration risk, governance and operational effort. Explain the evidence for a move and a limited migration/rollback plan. Familiarity with a product does not by itself justify replacement.
Avoid a shallow answer: “I deployed a model on a cloud.” Explain the controlled artifacts and operating responsibility.
DB-TQ14: How would you separate a customer implementation defect from a platform issue during an incident?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB007, DB033 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002; Tokyo FDE Manager CSQ427R97 / ATS 8535419002 for hands-on review. DBS01 official source, DBS04 official source.
What it may test — inference: Whether you make troubleshooting reproducible and useful to Engineering/Support; Manager review is labeled separately from IC work.
Evidence, numbers and decisions to include: Use an incident you may discuss. Show the symptom, impact, timeline, request/job identifiers, versions and smallest permitted reproduction. State the hypotheses tested and what was ruled out. Record actual duration or affected workload if shareable, and the temporary mitigation. Explain when you escalated, the evidence handed over and the customer communication owner.
Deep dive 1: You cannot reproduce the failure outside the customer environment. What can you provide?
Points to explain: Capture safe metadata, differences in data shape/configuration, relevant errors and reproducible synthetic conditions. State what the reproduction does not establish. Arrange approved diagnostic access rather than copying private customer data into a report.
Deep dive 2: A workaround restores service but leaves the cause uncertain. How do you close the incident?
Points to explain: Distinguish mitigation from resolution. Keep residual risk, the permanent-fix owner, monitoring and a follow-up verification step visible. For a Manager answer, explain how review and team learning improve the next incident without claiming every IC owns indefinite on-call.
Avoid a shallow answer: “I escalated it to Support.” Explain why, what evidence you supplied and what responsibility remained.
DB-TQ15: What technical contract would make a customer-specific component safe for another team to reuse?
Question status: Original inferred practice question; not an official interview question.
Grounding and scope: DB008, DB038, DB062, DB063 — Tokyo Sr. FDE CSQ327R52 / ATS 8568173002; Tokyo DSA CSQ427R266 / ATS 8813844002; global service scope for shared standards. DBS01 official source, DBS05 official source, DBS11 official source.
What it may test — inference: Whether a reusable asset has explicit boundaries and validation, rather than becoming a copied project with hidden customer assumptions.
Evidence, numbers and decisions to include: Choose a component you actually extracted or label the extraction a design exercise. Identify stable inputs/outputs, configurable differences, dependencies, permissions and failure behavior. Provide tests, supported versions, release responsibility and a migration note. Use observed reuse, setup time or defects if measured. Separate a reusable implementation from a product-roadmap request.
Deep dive 1: Only one customer has used it. How much abstraction is justified?
Points to explain: List the demonstrated common requirement and the assumptions still untested. Keep the interface small and customer-specific code explicit. Explain what a second use would teach before adding general options.
Deep dive 2: Another team changes it and breaks an older deployment. What must the asset contract cover?
Points to explain: Explain versioning, compatibility expectations, regression fixtures and change ownership. Provide an upgrade/rollback path and how consumers learn a breaking change. A shared repository alone is not a supported delivery contract.
Avoid a shallow answer: “We made a template.” State its inputs, limits, tests and maintenance owner.
An offline evidence worksheet
Select one production-path question, one failure-analysis question and one AI FDE question if it matches your target vacancy. Use a real sanitized project or a small synthetic case. This is an unexecuted preparation exercise, not a required assessment format.
| Record | What to write |
|---|---|
| Grounding | Question ID, fact IDs, exact role/requisition and source date checked |
| Case | Task, data revision, your responsibility and other owners |
| Decision | Baseline, alternatives, constraint and chosen design |
| Evidence | Trace, plan, test or measured result; units/window only where meaningful |
| Failure | What was observed, what was ruled out and what remains uncertain |
| Deep dive | Answer both follow-ups without inventing a result |
| Next step | Missing evidence, offline test and the limit it cannot establish |
Explain the case aloud, then ask someone to choose either follow-up. If you cannot explain a boundary, record the gap and revise the case. Keep a real achievement, a hypothetical design and an unexecuted test distinct. Continue with the delivery practice article for agreements, acceptance and customer ownership.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01