Scope and evidence

Source review
Cloud scope
AWS
Availability
Varies by feature
Exercise status
Exercises not run

Conditions and limits

  • AWS documentation scope; supported regions, model entitlements and workspace enablement must be checked individually.
  • Unity Gateway is GA, but service policies, unified trace table, external-model budget inclusion, managed memory/sessions and Agent Bricks CLI retain Beta boundaries.
  • Release rollout can take a week or more. No workspace, paid model call, latency, price or budget-enforcement result was tested.

Start with the job, then choose the AI role

A support representative asks, “Is order O-17 eligible for a return, and can you draft a reply?” A business analyst asks, “What share of last month's orders was returned?” A developer builds the support application. These are three different jobs. A conversational data interface, a workflow agent and a coding assistant need different inputs and permissions.

This lesson uses an original, synthetic support case. The application reads an order, retrieves the applicable policy, resumes the current conversation and drafts an answer. It may propose a return; recording that return requires a separately authorized business operation. Nothing here was run in a Databricks workspace.

Scroll horizontally to read the diagram.

Original role map: Genie serves business questions; a custom agent calls governed models and tools and must track executing identity; session memory, business state and MLflow evidence serve different purposes.

The figure is an original conceptual map. Arrows show selected application dependencies, not Databricks' internal deployment or automatic wiring. Open the diagram at full size. Read the lower row separately: a conversation record, a business transaction and an evaluation trace cannot substitute for one another. The tables below provide the text equivalent.

Genie has three distinct jobs

The current Genie documentation, revised 18 September 2026, names three experiences:

Product Who uses it and for what Support-case decision
Genie One Business users discover and interact with governed data assets through a simplified interface Ask about return trends or browse a dashboard
Genie Agents Data teams curate domain datasets, trusted metrics and business rules used in Genie One answers Define which orders and return definitions the question uses
Genie Code Developers and technical practitioners receive coding and data assistance inside the workspace Help author and inspect the support application's code

For the analyst, clarify whether the return rate counts orders or order lines and which date determines the month. A correct query over the wrong definition gives the wrong answer. For the developer, code assistance does not establish that generated code has the intended authorization or retry behavior. Genie Agents were formerly called Genie Spaces; use the current name when following present setup instructions.

Choose an agent only when the workflow needs it

The agent-building overview distinguishes Knowledge Assistant, Supervisor Agent and custom agents. A Knowledge Assistant suits assistance over domain knowledge. A Supervisor can coordinate Genie Agents, agent endpoints, Unity Catalog functions, MCP servers and custom agents. Custom Python agents allow application-specific behavior and integrate with MLflow tracing; Agent Bricks CLI is a Beta authoring/deployment route.

For the support case, begin with one agent and two conceptual tools: an order reader and a policy retriever. These are names for our design responsibilities, not SDK methods. Add a supervisor only if separately specialized components improve a defined acceptance test. More routing adds places where evidence, identity or a failure can be lost. Compare it against the single-agent baseline before accepting that complexity.

The history explains why evaluation belongs in the design. On 12 June 2024, Mosaic AI Agent Framework entered Public Preview with building, deployment, tracing and evaluation together. The 16 September 2026 release added managed memory and sessions in Beta; 29 September added Agent Bricks CLI in Beta. Managed state reduces a storage-operating task. It leaves the application's business rules and evaluation criteria to its author.

Follow the executing identity through each call

Unity Gateway governs AI traffic and access to registered services. Unity Catalog governs the assets behind those services. Gateway does not choose the support workflow or become its transaction ledger. Its controls apply to the traffic routed through the relevant governed services; an application's other data paths need their own checks.

Write an identity ledger before implementation:

Boundary Question to answer Concrete failure to test
User → application Which authenticated person and tenant made the request? Missing identity is rejected before state lookup
Agent → model service Which principal is presenting the credential? A principal without model access cannot query it
Agent → order reader Which caller's data privileges apply? A north-tenant user cannot retrieve a south-tenant order
Agent → state store Which principal can reach the store, and how is the actor selected? A model-supplied actor cannot redirect a lookup
Proposal → business write Which operation checks eligibility and records the result? A repeated request cannot create two returns

Do not assume the signed-in user executes every operation. The developer overview gives a specific AppKit example: agent model calls run as the app service principal; plugin tools on built-in agent routes run on behalf of the signed-in user. Standalone runAgent has no HTTP request context and runs model calls and tools as the service principal. These are AppKit-path semantics, not a promise for every framework or deployment.

For a custom model service, creation requires USE CATALOG, USE SCHEMA and CREATE SERVICE on the containing catalog/schema, plus privileges on destinations. Region support and Unity Catalog enablement are prerequisites; AWS GovCloud is excluded in the reviewed guide. Creating the service and allowing an application to query it are separate decisions. Inventory the actual model and endpoint available to the workspace instead of copying a model name from an old example.

Persist three kinds of state separately

Managed sessions store one interaction's JSON-compatible state, commonly ordered messages and tool results. Items are opaque to the service, and ordering is deterministic. When rebuilding a transcript with the documented client, session.list_items(order_by="create_time asc") requests chronological order; the default is newest first. A stored statement that a return succeeded is still only conversation content.

Managed memory preserves context across conversations. Its current retrieval contract is BM25 full-text ranking, with at most 100 top results, no pagination and no vector similarity search. The introductory phrase “semantic search” must not be expanded into an embedding-based guarantee.

State Synthetic content Lifetime and authority we choose
Session The O-17 question, policy tool result and draft reply Resume this case; do not treat the text as proof of a write
Memory “Prefers concise email replies” Reuse across cases only while relevant and permitted
Business record Return request ID, order ID, eligibility decision and committed status Application-owned schema and transaction/retry contract

Both managed offerings are Lakebase-backed and Beta. Their guides say the backing Lakebase instance is billed during preview, with no additional managed-state charge; pricing can change. Sharing a database technology does not establish an atomic transaction between state stores and an application's business record. The Lakebase lesson develops that write contract.

The store is the access boundary

Both guides require actor_id from trusted application context. An actor is a partition/grouping key, not an access-control boundary. Memory access is authorized at store level: a principal that can reach a memory store can read and write entries for every actor in it. The guide prescribes separate memory stores when strict tenant or user isolation is required. Session stores also authorize at store level; actor_id and metadata group or filter sessions without restricting access. Do not promise that an actor filter protects a hostile caller with store credentials.

For our design, use a trusted tenant/user key for ordinary lookup and review store grants separately. Determine whether the required isolation justifies separate stores or an alternative state service with the needed authorization model. Keep the per-store billing and management consequences visible. Deleting a session or session store does not delete memory; design a deletion procedure that inventories both stores and any business records independently.

Change a setting, predict the consequence

These choices are an unexecuted configuration worksheet. Example names are synthetic; they are not defaults. Retention durations, model prices, latency targets and account limits remain unknown until verified for the chosen environment.

Setting or control Meaning and scope Choice for the support design and expected effect
Memory actor_id and store grants Required actor key is chosen in trusted code; store access is the security boundary Bind the caller before exposing memory tools; test grants separately from actor filtering
Memory session_id, path, description Session association is optional; required path organizes entries; description assists retrieval Store only a durable preference at /preferences/reply-style.md; keep the case transcript in sessions
Session session_id Optional caller-chosen interaction ID; generated when omitted Use a synthetic case-o17 for a known case, then validate caller access before resuming it
Model service destination The governed service routes inference to configured models/providers Keep one permitted destination initially; changing a model requires rerunning the same cases
Budget action Send alert preserves access; Block usage prevents further requests in scope Choose blocking when continued use must stop; define what the application displays at exhaustion
Service-policy mode and phase Enforcing policies inspect ON CALL/ON RESULT; Log mode records without blocking Verify enforcement on a synthetic forbidden call before relying on a policy

Service policies are Beta and complement grants: grants answer who may call, policies decide how a particular interaction proceeds. ALLOW, DENY and ASK allow, block or hold for approval. A DENY can return HTTP 200 with databricks_service_policy details. Therefore HTTP success alone must not mark the business action completed. Enforcing evaluation errors deny; Log mode does not enforce. Model-based policy judges are non-deterministic and incur evaluator-model usage. These checks do not replace deterministic eligibility validation inside the return operation.

Gateway budgets apply to covered spend flowing through Gateway, exclude provisioned throughput and direct standalone Model Serving traffic, and optionally include external provider spend through a Beta setting. Blocking uses near-real-time estimates: in-flight requests continue, and the threshold is not an absolute final-invoice cap. External estimates can differ from the provider invoice. Budget tag filters match model-service tags, not request tags. In our worksheet, a support-project service tag scopes the workload; attaching the same tag only to requests would not create that budget scope.

Evaluate the answer, the action and the bill

The evaluation documentation describes evaluation and monitoring; the agent overview connects them with traces, human feedback, quality, cost and latency. A trace helps locate a bad policy retrieval, an unexpected principal or an expensive extra model call. It does not prove business correctness by itself.

Prepare a fixed synthetic fixture: O-17 belongs to tenant north; O-18 belongs to south; the policy permits returns within 30 days of delivery; the caller may read north orders and draft replies. The 30 days is our invented test rule. Record delivery and request dates explicitly so a human can determine eligibility. Include one contradictory document marked obsolete. Keep the documents, model identifier, instructions, tool permissions and dataset version fixed when comparing a one-agent and supervisor design.

Measure the fraction of answers with the correct order and policy, whether cited evidence supports the conclusion, unauthorized disclosure and mutation counts, end-to-end latency, model/evaluator tokens, retrieval calls and backing-state usage. Report cost per correctly completed permitted case with currency, billing date and exclusions once actual billing data exists. No latency or cost number was measured for this lesson. Score blocked and unresolved cases as those outcomes rather than treating every fluent response as completion.

Offline failure lab: specify the evidence before running

This is an unexecuted paper exercise. No credentials, deployed agent, live model call or customer data is needed. Draw at most seven boxes, label every tool call with its executing principal, and classify each persisted field as session, reusable memory or business state. Then complete this matrix:

Injected failure Expected application result Evidence the future test must retain
North user asks for O-18 Refuse the unauthorized read without disclosing the order Authenticated caller and denied tool result
Prompt requests another actor's memory Trusted actor binding cannot be overridden Actual actor passed by trusted code; separate store-grant test
Resume transcript newest first Detect the ordering error before using it as context Ordered item IDs and the reconstructed transcript
Delete the case but leave memory Report incomplete deletion, then follow the approved memory procedure Separate session and memory inventory/results
Policy returns HTTP 200 with DENY Display a blocked outcome; record no successful return Structured policy decision and business-record lookup
Same approved return is retried Resolve the existing request instead of a second mutation Application request ID and authoritative committed result
Evidence service times out or budget blocks Explain the unavailable evidence or exhausted budget Failed boundary and unresolved case status

Pass the design review when another reader can identify who may read each resource, which state survives a new conversation, which record proves a return, and what happens after a partial failure. Compare a simpler Genie answer, a document assistant and a custom workflow against those requirements. If a supervisor adds no testable benefit, keep the simpler design.

The Gateway release notes date Gateway GA to 4 August 2026 and its management API/developer tools GA to 16 September. They still identify service policies, agent services, the unified trace table and external-model budget inclusion as Beta. This lesson reviewed the rolling 4 September–4 October 2026 window; release notes warn that staged rollout may take a week or more. A GA label on the parent product cannot establish the state or account availability of each component.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Official documentationGenie Agents ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
02
Official documentationGenie ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
03
Official documentationUse agents on Databricks ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
04
Official documentationManaged agent memory ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
05
Official documentationManaged agent sessions ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
06
Official documentationUnity Gateway release notes ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
07
Official documentationUnity Gateway: developer overview ↗developers.databricks.comPublished: Unknown · Accessed: 2026-10-04
08
Official documentationAI governance with Unity Gateway ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
09
Official documentationService policies for AI securables ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
10
Official documentationManage budgets for Unity Gateway ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
11
Official documentationCustom model services ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
12
Official documentationEvaluate and improve ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
13
Official documentationDatabricks AWS product releases, September 2026 ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
14
Official documentationDatabricks AWS product releases, June 2024 ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.