A model API alone does not make a product
Calling a model API is enough for a first demo. A user-facing system also needs authentication, input limits, secrets, evidence-bearing data, failure presentation, observability, cost ceilings, and delivery. OpenAI, Anthropic, and Google provide models and developer APIs; ElevenLabs can provide voice components; Databricks emphasizes enterprise data and governance; Vercel and Cloudflare can host delivery, execution, and AI integration. This is a responsibility map, not a claim to describe every product from each company.
- 1User
- 2identity and input policy
- 3app server
- 4provider request
- 1Data
- 2permission-filtered retrieval
- 3evidence
- 4generated draft
- 1Tool request
- 2deterministic validation and approval
- 3side effect
- 1Trace and usage
- 2evaluation, budget, and incident response
Choose by the boundary a layer must own. Keep providers behind a server-side request boundary; browser code does not receive provider secrets. Normalize your own request, result, timeout, and error contract only after observing actual needs. A provider abstraction does not make model behavior, tool schema, price, retention, region, or quota identical. Confirm those on the current OpenAI, Anthropic, and Google documentation before use.
Evaluate a workflow, not a leaderboard
Make 20–50 representative cases from desired work and prohibited failures. For each, retain allowed evidence, expected output shape, whether a refusal is correct, and a deterministic or human scoring method. Run the same cases with fixed prompts, model IDs, temperature, maximum output, tools, and retrieval. Measure task success, unsupported claims, citation validity, schema failures, p95 latency, and cost. Keep a locked subset out of prompt and routing design.
Voice, enterprise data, edge execution, and web delivery add their own responsibilities. For speech, inspect ElevenLabs documentation for consent, transcript, and latency needs. For governed data, inspect Databricks generative-AI documentation and keep user authorization upstream of retrieval. For web runtimes, use the current AI SDK documentation and Workers AI documentation to implement a bounded request—not an unbounded autonomous call.
Start with one explicitly configured provider and one direct route. Do not implement fallback or automatic provider switching. A provider failure returns a stable error and stops dependent actions. Change the configured provider only after an observed need and evaluation of quality, data processing and cost. Provider selection is a recurring operational decision: store release date, access date, model revision, policy, test result, and rollback target together.
2024–2026 change: selection moved from a single call to a controlled path
The shift from chat completion demos in 2024 to tool-using and multimodal systems in 2025–2026 makes the provider boundary a product boundary. OpenAI's 2026-09-21 standards post and Google's 2026-09-15 language post are statements about their work and direction; neither supplies a portable quality, availability, or pricing guarantee.
Add a decision record before adding a second provider: which evaluated case fails, what input crosses the boundary, which identity and retention rule applies, how a model/tool version is pinned, and how a user-visible failure is returned. A routing rule should be deterministic and observable: a request either selects its explicitly configured route or fails with a stable error. Do not silently route confidential content to a different provider because a first call failed. Evaluate language separately from general task success: make fixtures for Japanese, English, mixed-script input, refusal, structured output, and the unsupported locale/region that must fail closed.
A September 20 r/LocalLLaMA discussion of FP4 inference is community experience, not a benchmark. It reinforces a practical control: measure the same task suite before and after a precision or runtime change instead of treating throughput claims as quality evidence.
OpenAI’s September 2024 o1 release made test-time compute an explicit product variable. A provider comparison can no longer use answer accuracy alone: hold prompt, tools, and retrieval fixed, then record latency and total request cost at each permitted reasoning setting. This is a measurement design, not a claim that later models share o1’s behavior.
MENTAL MODEL / VERIFICATION COST
The value of a decision depends on downstream work.
Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.
SOURCES
01YOUR NOTES