Scope and evidence
- Source review
- Cloud scope
- AWS
- Availability
- Not confirmed
- Exercise status
- Exercises not run
Conditions and limits
- Lakewatch launched in Private Preview on 2026-03-24. A later public release-stage transition and current enrollment were not verified on 2026-10-04.
- AWS is the scope of the synthetic cloud-log examples and companion documentation, not a claim of Lakewatch availability. Exact Lakewatch cloud/region eligibility and privileges remain unverified.
- The diagram is an original conceptual design. Detailed public Lakewatch APIs, field schemas, defaults and internal topology were not established by the reviewed sources.
- The event schema, thresholds, retention choices, detection workflow and response approval are author-created exercises. No workspace, connector, backtest or response action was executed.
Begin with a question about evidence
A successful sign-in appears from an unexpected country. Three minutes later, the same subject changes a cloud permission. Should the security team investigate, declare an incident, or contain an account? Each decision needs more evidence than the previous one. A country mismatch can result from a VPN; a permission change can be planned maintenance; an absent endpoint event can mean a broken collector.
Lakewatch is Databricks' security information and event management system, or SIEM. Its security lakehouse approach brings security telemetry and permitted business context together for detection and investigation. The March 24, 2026 announcement explicitly launched it in Private Preview. This is dated launch evidence. As of the October 4 review, a later public stage transition, current enrollment and an exact cloud/region matrix were not established. The broad availability language on the current product page does not settle those questions.
This lesson teaches an architecture decision and an offline investigation. It does not provide a verified Lakewatch setup procedure. AWS-oriented examples describe the learning scope, not Lakewatch availability in an AWS account.
Read the whole path before choosing a tool
Scroll horizontally to read the diagram.
The figure is an original teaching design, not Lakewatch's internal topology. Read the numbered stages left to right, then continue on the second row. The dashed boundary around the response decision is an approval design for this exercise, not an asserted product default.
The text equivalent is: collect → normalize → governed evidence → detect → investigate → response decision. Authorized business context joins the investigation; its outcome feeds future detection tests. A response executor needs a separate permission decision.
Think of security operations as an evidence desk. A witness statement, a consistent filing system, a lead and a decision to intervene serve different purposes. The analogy stops at trust: a normalized record can still contain incorrect source data, and an agent's fluent explanation cannot establish that a subject performed an action.
The launch article's architecture section describes OCSF, Lakeflow Connect, governed security data and detection as code. The product descriptions and FAQ describe agent-assisted triage, cross-source investigation and Unity Catalog governance. These are official capability descriptions. They do not specify every connector's support, normalization mapping or Lakewatch-specific privilege.
| Stage | Output and design responsibility |
|---|---|
| Collect | Preserve source provenance and delivery progress. The collector should have only the source access it requires; delayed delivery must remain visible. |
| Normalize | Align event time, subject identity and action meaning while retaining a reference to the raw event. Changing field names alone does not reconcile identities. |
| Govern evidence | Control who can read raw telemetry, normalized rows and enrichment data. Record the executing principal for each query. |
| Detect | Evaluate a versioned hypothesis and emit an alert with supporting event references. The rule does not establish an incident. |
| Investigate | An analyst and an authorized agent gather context, test alternatives and state missing evidence. Reading one table must not grant access to every business table. |
| Decide and respond | A designated owner decides whether the evidence warrants action. Record approval, executor, requested action and observed outcome separately. |
The identities and approval policy in this table are design requirements we chose. The reviewed public sources do not establish Lakewatch's internal agent identities or exact grant operations.
Four records answer different questions
Logs, traces, application state and detections can share storage technology while carrying different contracts. Keep their questions and retention policies separate.
| Record | What it can explain | What it does not establish by itself |
|---|---|---|
| Databricks audit event | A recorded platform action, its subject and response | Complete external identity/endpoint coverage or a confirmed compromise |
| MLflow agent trace | The agent's steps, inputs, outputs and timing | A complete security audit trail or authority to execute its proposal |
| Lakebase operational telemetry | Database sessions, waits, plans and resource behavior | A SIEM investigation across the enterprise |
| Agent session state | The history persisted for an application interaction | Independent source evidence, incident approval or successful containment |
In the AWS audit table reference, system.access.audit is Public Preview. Its documented fields include event_time, event_id, user_identity and response; identity_metadata can distinguish run_by and run_as. Most workspace logs are regional. A requester, an executing identity and an affected resource must therefore remain separate during an investigation. These are Databricks audit fields, not Lakewatch configuration fields.
MLflow tracing describes intermediate agent steps and quality investigation. A trace showing “revoke access” in a generated tool argument proves what the agent proposed; only the tool's independently checked result and follow-up state can support execution success.
The AWS system table reference lists a 365-day free retention period for audit logs and seven days for Lakebase observability. The audit reference documents an exception: deleting a workspace removes that workspace's events older than fourteen days from system.access.audit. The Lakebase telemetry reference labels that feature Beta and documents schema-level access across account projects. These are companion-table contracts, not Lakewatch retention defaults. A seven-day operational view cannot serve a thirty-day historical investigation without another appropriate evidence-retention path. Missing rows also require checking coverage and delivery, not assuming that nothing happened.
Managed agent sessions are Beta and use Lakebase for one interaction's durable state. The service persists opaque items; it does not add approvals or runs as first-class execution controls. Application conversation history and SIEM security evidence need separate ownership, access and deletion rules. This does not imply that Lakewatch internally uses those session APIs.
A synthetic case with provenance
All values and field names below are authored for the exercise. aws-audit-demo means an AWS-oriented cloud event fixture; it is not a real connector or CloudTrail event name. This compact schema is not an OCSF-compliant record or a Lakewatch API payload. Actual OCSF mappings require the applicable schema version, event class and connector documentation.
Save the following as events.json if you perform the offline exercise. Times and identities are synthetic. An upstream identity mapping has already established that employee-8 refers to the same subject within org-demo; a matching display name would not be sufficient.
[
{"event_id":"i-1","tenant":"org-demo","source":"identity-demo","occurred_at":"2026-10-01T09:00:00Z","received_at":"2026-10-01T09:00:05Z","actor":"employee-8","kind":"sign-in","outcome":"success","expected_country":"JP","observed_country":"GB"},
{"event_id":"e-1","tenant":"org-demo","source":"endpoint-demo","occurred_at":"2026-10-01T08:58:00Z","received_at":"2026-10-01T09:07:00Z","actor":"employee-8","kind":"device-status","outcome":"unknown"},
{"event_id":"c-1","tenant":"org-demo","source":"aws-audit-demo","occurred_at":"2026-10-01T09:03:00Z","received_at":"2026-10-01T09:04:00Z","actor":"employee-8","kind":"permission-change","outcome":"success","resource":"object-store-demo"},
{"event_id":"i-2","tenant":"org-demo","source":"identity-demo","occurred_at":"2026-10-01T09:12:00Z","received_at":"2026-10-01T09:12:05Z","actor":"employee-9","kind":"sign-in","outcome":"success","expected_country":"JP","observed_country":"JP"}
]Keep the raw source objects separately in a real design. A reference scoped by tenant, source and event ID lets an analyst inspect the original evidence and the normalization version. occurred_at answers when the source says the activity happened; received_at exposes collection delay. Joining on arrival order would put the endpoint event after the permission change despite its earlier event time.
Two competing explanations remain open: stolen credentials followed by a permission change, or an authorized operator using a VPN during approved maintenance. The next read-only investigation should compare the cloud request's executing identity with the relevant change approval and inspect the identity provider's session/MFA evidence. Asset ownership alone does not prove maintenance approval. The endpoint's unknown outcome supports neither explanation.
Write a detection hypothesis, then try to break it
Our hypothesis is: a successful country-mismatched sign-in followed within ten minutes by a successful permission change for the same tenant and canonical subject deserves review. Geography is a weak signal, so the output is an alert for investigation.
The following is original, unexecuted Python for a local file. It calls no Databricks service. window_minutes is an exercise parameter, not a documented Lakewatch setting. It uses event time, deduplicates source-scoped IDs and rejects conflicting duplicates or timezone-free timestamps.
import json
from datetime import datetime, timedelta
window_minutes = 10 # Author-chosen positive integer; minutes.
with open("events.json", encoding="utf-8") as file:
events = json.load(file)
def time_of(value):
result = datetime.fromisoformat(value.replace("Z", "+00:00"))
if result.tzinfo is None:
raise ValueError("Event time requires a timezone")
return result
unique = {}
for event in events:
if not event["actor"] or not event["tenant"]:
raise ValueError("Unresolved subject or tenant")
time_of(event["occurred_at"])
key = (event["tenant"], event["source"], event["event_id"])
if key in unique and unique[key] != event:
raise ValueError("Conflicting source event ID")
unique[key] = event
alerts = []
for signin in unique.values():
if signin["kind"] != "sign-in" or signin["outcome"] != "success":
continue
if signin["expected_country"] == signin["observed_country"]:
continue
for change in unique.values():
if change["kind"] != "permission-change" or change["outcome"] != "success":
continue
same_subject = (signin["tenant"], signin["actor"]) == (change["tenant"], change["actor"])
gap = time_of(change["occurred_at"]) - time_of(signin["occurred_at"])
if same_subject and timedelta(0) <= gap <= timedelta(minutes=window_minutes):
alerts.append({"rule_version":"exercise-v1", "evidence":[signin["event_id"], change["event_id"]]})
print(alerts)By inspection, the fixture should produce one alert referencing i-1 and c-1; that prediction has not been executed here. The code does not query endpoint evidence, create an incident, notify anyone or change a permission. Its nested scan is appropriate for this four-record fixture; a production implementation needs a measured query/stream design and a deliberate policy for late events and repeated evaluations.
Use the announced detection-as-code lifecycle as a design prompt: review the rule, test fixtures, backtest historical samples, assess alert quality, approve rollout and collect analyst feedback. The public description mentions YAML with SQL or Python and CI/CD; it does not supply a verified Lakewatch YAML schema or deployment command. Our Python is an independent learning example.
| Failure lab: change one input | Predicted result and question to answer |
|---|---|
Repeat the identical c-1 object |
One alert remains after deduplication. How would repeated rule runs avoid creating duplicate alerts? |
Move c-1 to 09:11 UTC |
No alert from this rule. A longer window catches more sequences but can correlate unrelated work. |
Change c-1 to another tenant |
No cross-tenant correlation. Test both correlation isolation and read authorization. |
| Remove the timezone or actor | Reject the input rather than guess. Where will the rejected record and collection error be visible? |
| Remove the endpoint event | The rule can still alert, but investigation evidence is incomplete. Mark the gap explicitly. |
| Add a verified maintenance approval | The alert still exists. An analyst may close it with evidence; approval is not retroactive deletion of the signal. |
Choose controls whose consequences you can explain
The following choices belong to our hypothetical design. They have no claimed Lakewatch defaults, supported bounds or configuration keys.
| Exercise choice | Purpose, tradeoff and verification |
|---|---|
| Ten-minute event-time window | Narrow the hypothesis. Compare one-, ten- and thirty-minute cases against labeled examples; measure noise and missed test cases separately. |
| Thirty days of raw evidence | Permit later reinspection in this exercise. Select real retention from investigation, privacy and contractual requirements; price storage, querying and any connector/model costs. |
| Asset owner and change approval as enrichment | Obtain useful context with a small data entitlement. A business-data reader needs its own grant; test a denied query and an unavailable approval record. |
| Human decision before containment | Keep read-only triage separate from mutation. A response needs an approved target, authorized executor and verified outcome; record failure or uncertainty explicitly. |
More retained data creates more investigative scope and more sensitive material to govern. Longer detection windows increase correlation opportunities and work. Broad agent permissions can remove manual steps while widening access exposure. Use collector freshness, normalization rejection rate, reviewed-alert relevance and analyst investigation time as separate measures. “No alerts” is not a success measure when a collector has stopped. Historical backtest accuracy also says little about a source absent from the sample.
Decide whether Lakewatch fits the job
A security lakehouse is a useful candidate when a team must correlate many sources, reuse governed business context and maintain repeatable detections. Before replacing an existing SIEM, establish current access eligibility and prove priority source coverage, evidence preservation, entitlement enforcement, alert quality and response integration against the team's actual process. Vendor cost and speed language is not a measured result or a contractual SLA for that workload.
An existing SIEM can be the better choice when its connectors and incident workflows already meet requirements and a new platform's current support remains unclear. For a narrow Databricks platform-access question, authorized audit-table analysis may be enough. For “why is this Postgres query slow?”, start with Lakebase operational telemetry. For agent behavior, use traces and application-state controls. Those narrower paths solve different questions and do not establish a complete SOC architecture.
Public sources reviewed on October 4 did not establish detailed Lakewatch API contracts, parser defaults, current entitlement rules or a per-region matrix. The product page's normalization link leads to general Lakeflow Connect material. This is a source-review limit, not proof that restricted documentation does not exist. Obtain those specifics before designing a real deployment; do not fill the gap with settings copied from another product.
Exercise: deliver an evidence map
Without using a live account, annotate the fixture and predict the failure-lab results. Produce a one-page investigation record containing the alert's rule version and evidence references, two competing explanations, one missing-data risk, the next read-only query, permitted data scope and the response decision owner.
Then add a separate proposed-response record with its target, approval, executor and outcome initially marked not executed. Explain what evidence would change that field. The exercise succeeds when it preserves provenance, avoids cross-tenant joins, treats correlation as a lead, and keeps alert → incident decision → executed response as three distinct states. This lesson's workflow, backtest and response lab remain unexecuted.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01