kumyu.Learn
← Kumyu home日本語 ↗

REGISTERED TAG

Agent tool use

Design tool contracts, permissions and action verification.

Published Updated
Intermediate → production

Cloudflare Sandbox SDK 1.0: execution state, snapshots, and migration boundaries

Read Sandbox SDK 1.0 as an execution-state contract: Durable Object control, immutable snapshots, explicit network policy, and a rehearsed one-way migration.

12 min↗
Published Updated
Intermediate → artifact and review design

Version agent work with Cloudflare Artifacts

Separate Git history, execution state, and large outputs; scope credentials and validate the exact commit before accepting agent work.

20 min↗
Published Updated
Intermediate → implementation

Design AI agents that survive disconnects

Separate a browser connection, durable work, and stored conversation state. Use PiHarness to reason about admission, reconnects, replay, and permission boundaries.

20 min↗
Published Updated
Beginner

Understand Skills by separating them from other mechanisms

Learn what changes and what stays the same, rather than memorizing names.

11 min↗
Published Updated
Intermediate → execution boundaries

Antigravity 09-2026: file edits, hook coverage, and stale-state rejection

A new file-tool contract changes both dispatch and authorization; rehearse one edit and reject stale state before migrating.

12 min↗
Published Updated
Beginner

Separate discovery, reading, and use

Understanding loading stages helps you isolate why a Skill is not working.

8 min↗
Published Updated
Beginner

Read catalogs by purpose rather than popularity

Separate official status, familiarity, and novelty from usefulness for your own work.

36 min↗
Published Updated
8 min↗
Published Updated
Beginner

Read safety, licenses, and dependencies

A Skill is text to read and can also lead to code and external connections.

8 min↗
Published Updated
Beginner

Create a small Skill: review a lesson against its sources

Reduce one failure of your own before adding many ready-made Skills.

12 min↗
Published Updated
Intermediate

Compare with and without a Skill fairly

Explain what your experiment changed instead of relying on one impression.

17 min↗
Published Updated
Intermediate

Read differences and retain them in an experiment ledger

Move from feeling that things improved to explaining what changed.

12 min↗
Published Updated
Beginner

What will Skills do as models improve?

Turn possible future value into testable questions instead of predictions stated as facts.

8 min↗
Published Updated
Intermediate → agent system design

Databricks AI roles: Genie, governed agents, memory and evaluation

Assign business questions, agent workflows and coding assistance to the right roles. Follow tool identities, state isolation, budget limits and failure evidence in a support case.

25 min↗
Published Updated
Intermediate

Jev Choice, Score, and Noul: do not confuse probability, confidence, and authority

Learn the roles of three question primitives and how to prevent concentrated distributions from becoming permission to automate.

15 min↗
Published Updated
Intermediate

Jev implementation lab: connect candidate generation, abstention, and E2E evaluation

Connect semantic judgment to deterministic candidate generation and safety boundaries, then design end-to-end evaluation with controls.

14 min↗
Published Updated
Intermediate → local serving evidence

Local-model runtime evidence: evaluate Strata without inheriting its claims

Read a dated local-runtime release as a configuration and evidence contract: code and weight rights, local serving, release assets, author measurements, and a bounded offline fixture.

11 min↗
Published Updated
Intermediate → decision evaluation

llama.cpp /v1/systemone: use a local decision endpoint as a routing signal

A dated local GGUF decision endpoint returns typed probabilities; it narrows routing, but never grants authority.

9 min↗
Published Updated
24 min↗
Published Updated
Evaluation design → offline repeat fixture

Agent evaluation: repeat the configured system before crediting the model

A bounded offline worksheet for measuring a specific agent configuration, its verifier authority, and repeat-to-repeat variation.

9 min↗
Published Updated
Beginner

Choose AI providers as operating layers, not a capability table

Design OpenAI, Anthropic, Google, ElevenLabs, Databricks, Vercel, and Cloudflare as model, data, execution, voice, and delivery layers.

7 min↗
Published Updated
Intermediate

Evaluation, safety, and cost design before placing generative AI in production

Connect evidence, authority, recovery, observability, and updates before measuring model cleverness, so answer quality remains reproducible in operation.

8 min↗
Published Updated
Intermediate

Data-connected AI with Databricks: Agent Bricks and Unity Gateway in practice

Design authorization, meaning, audit, and cost before retrieval quality when connecting AI to enterprise data.

8 min↗
Published Updated
Intermediate

AI apps on Vercel: put AI SDK, Gateway, and Workflow at product boundaries

Use Vercel’s agentic-infrastructure announcements to design streaming, tool calls, durable jobs, secrets, and evaluation as one web product.

8 min↗
Published Updated
Intermediate

Edge AI with Cloudflare: boundaries for Workers AI, AI Gateway, and Agents

Design state, authentication, cost, caching, and tool execution separately before shipping edge inference or an AI app quickly.

8 min↗