Connecting to data is not the same as being allowed to answer
Enterprise-AI failures often begin not with a missing model fact but with retrieval of information that must not be seen. The 2026-06-16 Agent Bricks announcement describes the need to consider token use, deployment, security, evaluation, monitoring, and context in addition to an agent loop. A model choice cannot make a production answer trustworthy while data revision, user rights, domain terms, cost, and failure tracing are unresolved.
- 1User
- 2identity
- 3authorized retrieval
- 4permitted evidence
- 1Evidence plus question
- 2model or agent
- 3cited draft
- 1Tool call
- 2policy and approval
- 3execute or deny
- 1Trace plus usage
- 2audit and evaluation
- 3improvement
Build a semantic contract first: a person defines which table, revision, aggregation, and authorization terms such as “revenue,” “customer,” and “active contract” mean. Before asking an LLM for SQL, expose only permitted views and a read-only query budget. Broad raw-table access makes model mistakes look like data-model mistakes.
The Unity Gateway announcement describes cross-model, agent, MCP, and skill controls with usage visibility, limits, and monitoring. Centralized tracking is useful, but the gateway must not decide whether a user may see a particular invoice. Carry the same authorization context from identity through data retrieval.
Start with one business question. Define an answer with its evidence table, revision, and time—not SQL output alone. Write target users, allowed views, forbidden columns, maximum rows, and query timeout in policy rather than only a prompt. Retain row count, truncation, time, and authorization in a trace; decide raw-log retention separately. Measure 20 fixed questions for correctness, appropriate holding when evidence is absent, absence of unauthorized material, p95 latency, and cost. Separate writes from a read-only version, showing target, diff, and approver before execution.
A concrete first evaluation
Choose a question such as “Which contracts renew this quarter?” and make the business owner define the current date boundary, contract status, source system, permitted fields, and acceptable citation. Create cases for a user who can see one organization, a user who can see none, a stale data revision, an ambiguous date phrase, and a request that asks for a prohibited column. The expected outcomes include an answer with evidence, an explicit hold, and an authorization denial; an answer without a source is not automatically better than a hold.
Run the cases with a fixed data snapshot. Inspect the retrieval or query plan, returned row count, model prompt size, generated citation, and final text. A high answer score paired with accidental over-disclosure is a failed run. When source schemas or semantic definitions change, version the contract and re-run the same suite before claiming that an agent was improved. Governance is the ability to explain which user saw which revision of which evidence, while controlling what a model may do next.
2024–2026 change: governed context is now evaluated as a decision surface
The 2024 RAG question was often “can the model find a document?” In 2025–2026 the hard question is whether a particular principal may use a particular revision to take a defined action. Databricks introduced ai_decide on 2026-09-30 as a beta decision-oriented AI Function. Those are vendor claims and release signals, not proof of correctness, lower cost, or eligibility for any workspace.
Keep generation and control decisions separate. A deterministic policy gate must decide whether a principal can read a table or invoke an action. A probabilistic classifier may only propose a queue, score, or review state; record its version, threshold, calibration data, and the human consequence of each false positive and false negative. For a SQL-scale trial, use a frozen, synthetic dataset; run deny-by-policy cases before happy paths; and retain only query identifiers, decision labels, and approved audit fields. A low-latency decision is unsafe if its inputs include a column the requester should never have seen.
A September 24 r/selfhosted discussion about agents using self-hosted documentation is practitioner experience, not evidence about governed-data behavior. It supports the exercise of tracing retrieval paths on synthetic data; it cannot validate authorization, pricing, or results in a Databricks workspace.
Databricks introduced DBRX in March 2024 as an open general-purpose model and reported its own benchmark and efficiency results. That era made model access and serving central concerns; the governed-agent question adds a separate boundary: which principal may retrieve which data revision. Reproduce one DBRX-era-style answer fixture with a permitted view and the same question with a forbidden column, then require an explicit denial in the latter case.
MENTAL MODEL / VERIFICATION COST
The value of a decision depends on downstream work.
Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.
SOURCES
01YOUR NOTES