Work from the vacancy, then practice a decision
These are original inferred practice questions, not questions Anthropic has announced it asks. They connect specific duties to evidence you can prepare; their purposes are editorial reasoning, not disclosed hiring criteria. Fact IDs refer to the verified role and hiring reference, checked 2026-10-05. No internal interview loop or pass threshold is assumed.
Tokyo Applied AI Engineer is the main implementation reference. Tokyo Architect is explicitly pre-sales. London FDE and Sydney Engineer questions carry their own labels; a London MCP responsibility is not a confirmed Tokyo interview task. Use projects you actually contributed to. If you have only a prototype or an offline exercise, identify that limit instead of adopting an invented production success.
The candidate AI-use policy displays a last-update date of 2025-07-10 (A010–A014). AI may help preparation and refinement; take-home and live work exclude AI unless explicitly allowed. Rehearse an unaided version and ask what each assessment permits. Ordinary lookup permission does not change that boundary.
Input contracts, evaluation, recovery and stakeholder decisions are shared practices across companies. The role facts determine which responsibility to emphasize; these exercises do not invent proprietary Anthropic methods.
- 1Selected vacancy and fact IDs
- 2actual project evidence
- 3decision and rejected alternative
- 1Decision
- 2evaluation and failure case
- 3implementation and recovery
- 4ownership and handoff
- 1Unaided rehearsal
- 2deep-dive challenge
- 3revise evidence gaps
Original preparation flow. Arrows describe the candidate's work, not Anthropic's recruiting rounds. Technical depth means keeping input, authority, measurement and failure behavior connected. When a follow-up changes a constraint, explain which decision changes and which evidence is still missing.
AN-TQ01: Where would you place a Claude pilot in a manual triage workflow?
Scope and evidence: Tokyo Applied AI Engineer; Japan enterprise pilot — A006, A019, A020, A025; All roles; technical roles where stated, Applied AI Engineer.
Inferred practice question: For a Japanese enterprise that manually classifies incoming requests, how would you choose the first Claude-assisted step and its architecture?
What it may test — editorial inference: Whether you can turn an ambiguous customer need into a small, falsifiable implementation rather than selecting a complex agent architecture first.
Evidence and decisions to explain: Use a real workflow you observed: incoming fields, operator decisions, downstream system, volume and cost of a wrong classification. Draw the smallest interface that accepts an authorized input and returns a reviewable recommendation. State the non-AI baseline, the exact user decision being improved and a stopping condition. Explain who agreed to the pilot boundary and which integration you postponed. Bring actual counts and dates if recorded; label any constructed example as an exercise.
Deep dives:
- The sponsor wants an autonomous end-to-end agent immediately. What do you do? Separate the desired outcome from the requested implementation. Identify the highest-consequence action, compare a recommendation-only pilot with automation and agree on evidence required to expand authority. Record the sponsor decision and unresolved dependencies.
- The pilot succeeds only for one department. Is it ready to expand? Compare input distributions, operators and failure costs in the next department. Show which assumptions held in the first cohort, what must be retested and the rollout decision owner. One successful department is evidence for that scope, not universal readiness.
Avoid a shallow answer: “We would build an agent and optimize the prompt.” This omits the workflow, baseline, authority and testable decision.
AN-TQ02: What belongs in the adoption evaluation suite?
Scope and evidence: Tokyo Engineer or Architect; customer-specific evaluation — A018, A020, A030; Applied AI Engineer, Applied AI Architect.
Inferred practice question: A pilot appears useful in a demonstration. How would you build an evaluation suite that supports a customer adoption decision?
What it may test — editorial inference: Whether you distinguish a convincing demonstration from evidence covering the customer workflow and its failure costs.
Evidence and decisions to explain: Describe a real evaluation set: sampling period, source of cases, permitted data, expected outcomes, reviewers and disagreement handling. Separate frequent cases from rare consequential failures. Compare with the current workflow and report counts as well as percentages, with a denominator and confidence limits where justified. Explain the acceptance decision and its owner rather than presenting an aggregate score as permission to deploy.
Deep dives:
- The customer asks for a single accuracy number. What do you report? Offer the overall number with its denominator, then the critical slices and unacceptable outcomes. Explain what the number excludes, who reviewed ambiguous cases and why a high average cannot waive an agreed failure constraint.
- The suite was built from demonstrations selected by the sales team. What changes? Keep those demonstrations as a labeled set, collect representative and adversarial cases separately and prevent tuning against the final holdout. Record selection bias and the adoption decision that must wait for broader evidence.
Avoid a shallow answer: “Accuracy improved.” Without case selection, denominator, failure categories and a decision rule, this does not establish customer readiness.
AN-TQ03: How do you isolate an integration blocker?
Scope and evidence: Tokyo Applied AI Engineer; integration into existing enterprise infrastructure — A021, A025; Applied AI Engineer.
Inferred practice question: Claude works in a prototype but fails when connected to the customer’s existing system. How do you diagnose and resolve the blocker?
What it may test — editorial inference: Whether you can identify the failing boundary and work with its owner instead of attributing every failure to the model.
Evidence and decisions to explain: Bring an actual interface diagram and one redacted failing trace. Separate input/schema problems, identity and permission checks, upstream availability, model behavior and downstream writes. Explain the reproduction, smallest isolating test, responsible owner and repair verification. Quantify affected requests and period when available. Preserve customer authorization: reproducing a failure does not give permission to broaden access or copy production data.
Deep dives:
- You cannot obtain customer production logs. How do you progress? Ask for a permitted redacted sample or customer-run diagnostic; construct a synthetic reproduction with matching shape. State which conclusion the fixture supports and which production condition remains unverified. Do not present the fixture as production proof.
- A workaround needs a broader service-account permission. Would you use it? List the exact operation and resource missing, compare a narrow permission change with a design change and involve the customer’s authority owner. Verify the intended operation and a denied out-of-scope case; record whether the workaround is temporary.
Avoid a shallow answer: “We improved the prompt.” A boundary failure needs evidence identifying where the request stopped.
AN-TQ04: Can you explain a small Python implementation and its failures?
Scope and evidence: Tokyo Engineer Python preparation; general technical-interview logistics, not a confirmed task — A002, A003, A012, A013, A022; All roles; technical roles where stated, All candidates, Applied AI Engineer.
Inferred practice question: As an offline exercise, implement a transformation from customer records to validated results and explain how invalid records and duplicate inputs are handled.
What it may test — editorial inference: Whether Python fluency includes an explicit data contract, understandable code and reasoning about edge cases. This is an invented practice exercise.
Evidence and decisions to explain: Define input types, allowed fields, validation errors and output behavior before coding. Show a small implementation using familiar standard-library tools, tests for a valid record, missing field, unexpected type and duplicate. Explain where deduplication belongs and whether partial success is acceptable. Rehearse aloud without AI; permitted lookup is different from AI permission. Colab/CodeSignal are examples on the general page, not a promise of this exercise.
Deep dives:
- The customer asks to preserve valid records when one fails. What changes? Specify per-record status and error representation, deterministic ordering and the caller’s retry behavior. Demonstrate that failed records remain visible and are not silently dropped. Explain why the original all-or-nothing contract was changed.
- The input no longer fits in memory. How would you revise it? Identify the measured size and required ordering, then compare streaming, bounded batches and external state for deduplication. Explain the memory/time trade-off and recovery behavior; do not claim a scalable solution without exercising the relevant volume.
Avoid a shallow answer: Naming libraries without describing input, errors, tests or complexity. Also avoid using AI during an assessment merely because lookup is allowed.
AN-TQ05: When would you stop changing the prompt?
Scope and evidence: Tokyo Engineer prompting; London FDE advanced prompting background — A006, A019, A048; All roles; technical roles where stated, Applied AI Engineer, Forward Deployed Engineer.
Inferred practice question: A customer solution fails on a recurring class of requests. How do you decide between prompt changes, better input handling and a different system design?
What it may test — editorial inference: Whether you use experiments to locate the cause instead of treating prompting as an unlimited repair mechanism.
Evidence and decisions to explain: Choose a failure you actually investigated. Show the input, intended result, failure category and a fixed set of comparison cases. Change one variable at a time and track improved cases, regressions, latency and effort. Explain whether the failure came from missing evidence, an ambiguous task, an impossible requirement or execution outside the model. State the evidence that made you stop prompting and who accepted the alternative.
Deep dives:
- A longer prompt improves the demo but hurts other cases. How do you decide? Compare the same cases, record which instruction caused the change and report the regression slices. Consider clearer task separation or deterministic validation. Agree which trade-off is acceptable rather than hiding the failing examples.
- The model lacks current customer facts. Can instructions solve that? Separate absence of evidence from following instructions. Identify an authorized source and a way to return “insufficient evidence”; compare retrieval or a customer lookup with prompt-only changes. These are proposed design options, not declared Anthropic product defaults.
Avoid a shallow answer: “We iterated until it worked.” Name the controlled comparison and the failure that prompting could not resolve.
AN-TQ06: How would you release a changed model or prompt?
Scope and evidence: Tokyo Applied AI Engineer; recent production-LLM experience is preferred — A021, A024; Applied AI Engineer.
Inferred practice question: A production customer workflow needs a model or prompt update. How would you decide whether to release it and recover if it regresses?
What it may test — editorial inference: Whether production experience includes change control and observable recovery, rather than one successful deployment.
Evidence and decisions to explain: Describe an actual change with versioned inputs, prompts, configuration and evaluation cases. State the prior release, reason for change, measured quality/cost/latency differences and agreed release gates. Explain a bounded rollout, monitoring owner, rollback target and conditions that stop expansion. If the change was never shipped, distinguish local validation from production evidence.
Deep dives:
- Offline evaluation improves but users complain after rollout. What do you inspect? Compare the production population with the evaluation set, changed workflow and timing. Inspect permitted traces for a representative complaint, separate quality from usability and keep the previous configuration available. Add the missing slice before expanding.
- Rollback restores quality but loses a new feature. Who decides? Present the incident impact, affected cohort, benefit of the new feature and remaining alternatives to the accountable customer owner. Explain the temporary operating mode and communications. A reversible release still needs an explicit decision.
Avoid a shallow answer: “We always use the newest model.” Release evidence must be tied to the customer workload.
AN-TQ07: How do you design for a burst in customer demand?
Scope and evidence: London FDE; production applications and deployment at scale — A043, A048; Forward Deployed Engineer.
Inferred practice question: An embedded customer application works at pilot volume but misses latency and budget goals during bursts. How would you investigate and redesign it?
What it may test — editorial inference: Whether you connect deployment scale to workload shape, resource limits and the user-visible outcome.
Evidence and decisions to explain: Use measured workload facts: peak arrival rate, concurrency, input/output size, latency percentiles, error counts and cost unit for the actual environment. Break the request path into waiting, model work, tools and downstream services. Compare bounded admission, fewer calls, asynchronous work or cacheable results where appropriate. Explain what each alternative gives up and test a representative load before claiming capacity.
Deep dives:
- The average latency is acceptable but a small cohort waits much longer. What matters? Report the tail, affected request class and user consequence. Locate queueing or slow dependencies for that cohort. Choose a cohort-specific limit or path rather than celebrating an average that conceals the failure.
- Requests exceed the agreed capacity. What should users see? Define an explicit queued, rejected or reduced-scope result with a bounded wait and recovery path. Preserve work and prevent repeated side effects. State the customer owner’s acceptable degradation; do not silently pretend every request completed.
Avoid a shallow answer: “We add more infrastructure.” Show the actual bottleneck and resulting quality, cost or responsiveness trade-off.
AN-TQ08: What is the boundary of a customer MCP server?
Scope and evidence: London FDE; production MCP artifacts, not a published interview task — A043, A044, A045; Forward Deployed Engineer.
Inferred practice question: For a customer agent that needs business-system actions, how would you design and verify the boundary of an MCP server?
What it may test — editorial inference: Whether a production artifact has a bounded contract and authority rather than merely exposing a collection of tools.
Evidence and decisions to explain: Describe a real or explicitly hypothetical action: caller identity, allowed resource, input schema, output, error and side effect. Separate reads from writes, bind access to customer authorization and record where a human decision is required. Explain validation, timeouts, deduplication and redacted observation. Show a permitted action and a forbidden cross-resource action. These are original design considerations, not asserted Anthropic implementation defaults.
Deep dives:
- A call times out after a possible write. What happens next? Identify the observable operation ID and whether completion can be queried safely. Distinguish unknown from failed; avoid repeating a side effect without deduplication or reconciliation. Explain what the user sees and who resolves uncertainty.
- Retrieved content tells the agent to invoke an unrelated action. How do you respond? Treat source content as data and preserve the intended task and authorization boundary. Verify the action independently at the tool boundary, show the denied test case and retain a safe result without exposing the malicious content to privileged control.
Avoid a shallow answer: “MCP connects the agent to everything.” Enumerate the allowed contract and the actions that remain denied.
AN-TQ09: When is a sub-agent worth the coordination cost?
Scope and evidence: London FDE; production sub-agent artifacts — A006, A045, A048; All roles; technical roles where stated, Forward Deployed Engineer.
Inferred practice question: A customer workflow contains research, drafting and approval steps. When would you split work among sub-agents, and when would you keep a simpler path?
What it may test — editorial inference: Whether decomposition follows independent work and measurable benefit, with ownership of the final result.
Evidence and decisions to explain: Start with a simpler baseline and draw each task’s inputs, permissions, output contract and dependencies. Explain which steps can proceed independently and which must wait for another result or an authorized decision. Compare quality, elapsed time, calls and coordination failures on the same workload. Name the component responsible for checking inconsistent results and terminating work. Do not claim that more agents imply better performance.
Deep dives:
- Two sub-agents return contradictory recommendations. What resolves it? Keep their evidence and uncertainty separate. Specify a deterministic consistency check, an evidence review or a human decision appropriate to the workflow. A third generated opinion is not automatically stronger evidence.
- One sub-agent fails late. Must the whole workflow restart? Identify the durable outputs that remain valid, the dependency that failed and a bounded retry or explicit partial result. Avoid redoing completed writes. Describe when stale evidence invalidates otherwise reusable intermediate work.
Avoid a shallow answer: “Parallel agents make it faster.” Show the baseline, coordination cost and responsibility for disagreement.
AN-TQ10: How would you package a reusable agent skill?
Scope and evidence: London FDE agent skills; US Engineer reusable knowledge assets — A045, A046, A061; Forward Deployed Engineer, Applied AI Engineer, Enterprise Tech.
Inferred practice question: A useful customer procedure should become a reusable agent skill or comparable execution asset. What would you package and how would you verify its limits?
What it may test — editorial inference: Whether reusable guidance includes a usable contract, evidence and maintenance rather than a copied successful prompt.
Evidence and decisions to explain: Use one procedure you may describe. Record its intended user, trigger, input/output, required environment, authorized operations, failure handling and owner. Separate reusable guidance and tests from customer-specific identifiers and secrets. Show cases where the asset should run, decline or ask for information. Version it against dependencies and record a change review. No particular skill file format or internal Anthropic repository is assumed.
Deep dives:
- Another customer has different approval rules. Can the same asset be used? Identify which policy is an explicit input and which behavior must remain fixed. Require the new customer owner to approve the adapted contract and rerun relevant cases. Reuse of text is not proof of reuse of authority.
- The asset is reused frequently but its failure rate is unknown. Is it successful? Separate adoption from correctness. Define a permitted outcome sample, failure categories, maintenance responsibility and a withdrawal condition. Include the effort saved only when measured; downloads or invocations alone do not establish quality.
Avoid a shallow answer: “We saved the prompt as a template.” Include scope, permission, failure tests and an update owner.
AN-TQ11: What is an unacceptable failure in this workflow?
Scope and evidence: Tokyo Engineer and London FDE; reliability/safety responsibilities — A007, A018, A044, A048; All roles; technical roles where stated, Applied AI Engineer, Forward Deployed Engineer.
Inferred practice question: How would you define and test the failures that prevent a customer AI workflow from entering production even when its average quality is good?
What it may test — editorial inference: Whether safety and reliability become observable release constraints tied to actual consequences. This is not a disclosed company rubric.
Evidence and decisions to explain: Identify affected people, data and actions, then classify failures by consequence: wrong recommendation, unauthorized disclosure, duplicate write or inability to recover. Specify what each test observes and the safe terminal result expected. Separate a model’s suggestion from the application’s permission check and human decision. Describe the release owner and evidence needed to waive or remove a blocker; do not claim every risk can be reduced to one accuracy score.
Deep dives:
- The customer accepts the risk verbally to meet the deadline. What do you do? Clarify the specific consequence, authority of the decision-maker and applicable customer requirements. Offer a narrower release or review gate and record the unresolved condition. Escalate beyond your authority rather than treating urgency as authorization.
- A rare harmful case cannot be reliably reproduced. Can testing end? Preserve the observed evidence, identify the uncertain mechanism and add bounded probes or conservative controls. State what remains unknown and the monitoring/stop plan. Absence in a small sample is not evidence that the failure is impossible.
Avoid a shallow answer: “We follow safety best practices.” Explain the prohibited outcome, test observation and person who can decide.
AN-TQ12: What happens when the workflow fails after go-live?
Scope and evidence: Tokyo Engineer post-go-live advice; London FDE production delivery — A021, A043, A044; Applied AI Engineer, Forward Deployed Engineer.
Inferred practice question: A customer workflow stops midway after deployment. How do you help restore service and preserve a correct account of completed work?
What it may test — editorial inference: Whether operational support includes state, ownership and safe recovery instead of simply rerunning the request.
Evidence and decisions to explain: Bring an incident you actually handled, or label a tabletop scenario. Reconstruct intent, completed stages, uncertain side effects and affected users from permitted evidence. Explain the mitigation, customer incident owner, retry limits and reconciliation. Distinguish restored service from repaired data and from a permanent fix. Report actual duration and affected counts only if known, then add the failure to the regression set.
Deep dives:
- The upstream call failed but the downstream system may have changed. Do you retry? Identify the downstream operation and a safe way to inspect completion. Keep uncertain work visibly unresolved until reconciled. Use idempotency or deduplication where the design supports it; do not assume an upstream error proves no write occurred.
- The quick mitigation works but requires manual reviews. How do you hand it over? Document the temporary procedure, queue owner, capacity and escalation/expiry conditions. Train the operator, verify one representative recovery and agree the permanent-fix priority. Mitigation is a supported operating mode, not disappearance of the incident.
Avoid a shallow answer: “We restarted the service and solved it.” Separate recovery, data correctness, cause and follow-up ownership.
AN-TQ13: How does a code-review workshop change the customer’s implementation?
Scope and evidence: Sydney Applied AI Engineer; workshops and code reviews, not a confirmed Tokyo duty — A054, A055, A056; Applied AI Engineer.
Inferred practice question: Design a hands-on workshop for a customer engineering team whose prototype needs production evaluation and integration improvements.
What it may test — editorial inference: Whether teaching produces a verified engineering change and transfers capability, rather than a demonstration the customer cannot maintain.
Evidence and decisions to explain: Use a real workshop or a labeled plan. Identify participants’ starting skills, the code boundary, one observable learning goal and permitted data. Plan a small change, code review and a test the customer runs independently. Show what artifact remains, who maintains it and how understanding was assessed. If citing outcomes, report attendance separately from correct independent execution.
Deep dives:
- The most senior attendee understands but the operators do not. Is enablement complete? Assess the roles that will run and debug the system. Adapt the exercise, documentation and ownership to those operators, then verify their independent handling of a failure. Senior approval is not proof of operational capability.
- The customer wants you to fix everything during the workshop. What do you choose? Separate an urgent blocker from the agreed learning objective. Prioritize one transferable pattern, record follow-up work and make responsibility explicit. Explain the trade-off between immediate repair and the customer’s ability to maintain the next change.
Avoid a shallow answer: “We gave training.” State what the customer could do afterward and how that was observed.
AN-TQ14: What should become a repeatable deployment pattern?
Scope and evidence: London FDE and US Enterprise Tech Engineer; reusable deployment knowledge — A046, A061; Forward Deployed Engineer, Applied AI Engineer, Enterprise Tech.
Inferred practice question: After a successful customer implementation, how do you decide what becomes a reusable technical pattern and what remains customer-specific?
What it may test — editorial inference: Whether generalization preserves assumptions, evaluation and maintenance rather than exporting private code or overbuilding a framework.
Evidence and decisions to explain: Choose an artifact you were authorized to reuse. Separate domain assumptions, integrations, configuration, evaluation cases and customer-sensitive material. Describe the second use case that tests generality, required dependencies and failure modes. Compare a documented recipe with a library or template, including support burden. Show measured reuse only when recorded and name the maintenance owner and withdrawal condition.
Deep dives:
- Only one customer has used it. Should it become a general framework? Record it as a scoped recipe with explicit assumptions. Identify what evidence from another use case would justify abstraction, and avoid claiming generality from a single success. A smaller artifact may be easier to maintain.
- The second customer needs incompatible behavior. What does that mean? Check whether the difference is configuration, a genuinely different domain constraint or a violated assumption. Keep separate patterns when the common abstraction obscures correctness. Explain the evidence behind the split and update the usage boundary.
Avoid a shallow answer: “We generalized everything.” State what you deliberately did not generalize and why.
AN-TQ15: What do you discover before recommending an adoption architecture?
Scope and evidence: Tokyo Applied AI Architect; explicitly pre-sales, API and Claude for Work — A028, A029, A030, A033; Applied AI Architect.
Inferred practice question: A company is considering both Claude API use and Claude for Work. How would you discover which needs require custom integration and which need a different adoption path?
What it may test — editorial inference: Whether pre-sales technical advice begins with users, systems and evidence rather than assuming that every customer needs custom code.
Evidence and decisions to explain: Map the users, tasks, systems, data boundaries, administrator, buyer and intended outcome. Distinguish an employee workflow from a customer-product integration, then list the product behavior and control requirements that must be confirmed in current documentation or with the appropriate team. Propose a use-case evaluation and implementation owner. Do not invent feature parity, licensing, data terms or deployment capabilities from the job description.
Deep dives:
- The executive wants one product for every use case. What do you show? Show the distinct workflow requirements, integration and control dependencies, plus the evidence still needed for each option. Explain cost/complexity as measured or clearly estimated, and propose a small evaluation rather than promising one product solves all cases.
- A pre-sales recommendation reaches an implementation detail outside your authority. What next? Bring the documented requirement and unresolved technical question to the Engineer or relevant product owner. Preserve ownership, assumptions and customer expectations in the handoff. Technical credibility includes knowing which claim still needs verification.
Avoid a shallow answer: “API for developers, Work for everyone else.” Users alone do not establish the integration, control or value requirement.
Offline exercise: make one technical answer falsifiable
Choose one question whose scope matches your vacancy. On one page, write the input/output boundary, your contribution, baseline, evaluation denominator and period, failed cases, chosen alternative and release/stop decision. Add one artifact you are permitted to describe: a redacted diagram, test outline or decision record. Do not expose customer secrets.
Rehearse the main answer in three minutes without AI, then let a reviewer choose either deep dive and change one constraint. This is an original, unexecuted practice method; three minutes is a rehearsal setting, not Anthropic's interview duration. Record what you could not substantiate. Use production AI controls to review evaluation and authority, then practice customer delivery without transferring role-specific ownership.
MENTAL MODEL / VERIFICATION COST
The value of a decision depends on downstream work.
Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01