kumyu.Learn
← Kumyu home日本語 ↗

REGISTERED TAG

Evaluation

Define conditions and criteria; distinguish estimates from measured outcomes.

Published Updated
Intermediate → artifact and review design

Version agent work with Cloudflare Artifacts

Separate Git history, execution state, and large outputs; scope credentials and validate the exact commit before accepting agent work.

20 min↗
Published Updated
Intermediate → implementation

Design AI agents that survive disconnects

Separate a browser connection, durable work, and stored conversation state. Use PiHarness to reason about admission, reconnects, replay, and permission boundaries.

20 min↗
Published Updated
Authentication → offline boundary fixture

Access service authentication: verify the machine principal on every request

Use an offline request matrix to separate a service credential, a browser session, an edge decision, and an authorized application action.

13 min↗
Published Updated
Beginner

Computer use: build systems that can observe, act, and recover

A practical map of browser and desktop automation: perception, planning, action, verification, and safe recovery.

12 min↗
Published Updated
Beginner

Understand Skills by separating them from other mechanisms

Learn what changes and what stays the same, rather than memorizing names.

11 min↗
Published Updated
Intermediate → execution boundaries

Antigravity 09-2026: file edits, hook coverage, and stale-state rejection

A new file-tool contract changes both dispatch and authorization; rehearse one edit and reject stale state before migrating.

12 min↗
Published Updated
Beginner

Separate discovery, reading, and use

Understanding loading stages helps you isolate why a Skill is not working.

8 min↗
Published Updated
8 min↗
Published Updated
Beginner

Read safety, licenses, and dependencies

A Skill is text to read and can also lead to code and external connections.

8 min↗
Published Updated
Beginner

Create a small Skill: review a lesson against its sources

Reduce one failure of your own before adding many ready-made Skills.

12 min↗
Published Updated
Intermediate

Compare with and without a Skill fairly

Explain what your experiment changed instead of relying on one impression.

17 min↗
Published Updated
Intermediate

Read differences and retain them in an experiment ledger

Move from feeling that things improved to explaining what changed.

12 min↗
Published Updated
Beginner

Maintain comparisons across updates

Keep Skills in maintainable units rather than simply adding more.

8 min↗
Published Updated
Beginner

What will Skills do as models improve?

Turn possible future value into testable questions instead of predictions stated as facts.

8 min↗
Published Updated
Intermediate → agent system design

Databricks AI roles: Genie, governed agents, memory and evaluation

Assign business questions, agent workflows and coding assistance to the right roles. Follow tool identities, state isolation, budget limits and failure evidence in a support case.

25 min↗
Published Updated
Beginner

Hacker News observation: turn attention into bounded technical experiments

A dated HN API snapshot, read as a set of technical hypotheses rather than a popularity ranking.

13 min↗
Published Updated
Introduction → production planning

A map of creative deliverables and work

Distinguish a single image from an interactive service.

24 min↗
Published Updated
Introduction → production planning

Use AI to propose candidates under constraints

Separate the model, inputs, generation, editing and verification.

10 min↗
Published Updated
Introduction → production planning

Turn visual direction into words and specifications

Give instructions through composition, color, light, shape and purpose.

8 min↗
Published Updated
Intermediate

Practical image workshop: Creating a single advertising visual using a verifiable process

An exercise in creating still images containing products and people by dividing them into specifications, references, composition, creation, local corrections, typesetting, and release checks.

14 min↗
Published Updated
Introduction → animation and video workflows

Produce motion and video through separate workflows

Work with timing, states, editing and sound

9 min↗
Published Updated
Introduction → reproducible production comparisons

Measure quality and cost to guide the next step

Keep tokens, images, seconds and credits distinct

23 min↗
Published Updated
Beginner

Jev foundations: typed decisions outside text generation

Use Jev/System One as a small typed decision component with candidates, state, and probabilities alongside ordinary code.

13 min↗
Published Updated
Intermediate

Jev Choice, Score, and Noul: do not confuse probability, confidence, and authority

Learn the roles of three question primitives and how to prevent concentrated distributions from becoming permission to automate.

15 min↗
Published Updated
Intermediate

Jev implementation lab: connect candidate generation, abstention, and E2E evaluation

Connect semantic judgment to deterministic candidate generation and safety boundaries, then design end-to-end evaluation with controls.

14 min↗
Published Updated
Beginner

Map AI, machine learning and LLMs

Match the problem you want to solve to a mechanism, rather than treating AI as one uniform tool.

11 min↗
Published Updated
Beginner

Quantization, KV cache, and runtimes for a local LLM that stays responsive

Separate model weights, working memory, and generation speed; choose MLX, llama.cpp, or vLLM and turn the choice into a reproducible local evaluation.

11 min↗
Published Updated
Beginner

Understand models, tokens and inference

Follow the transformation from text to numbers, then from numbers to the next token.

8 min↗
Published Updated
Beginner

Choosing MiniMax and local LLMs: separate open weights, APIs, and feasibility

A practical research and experiment guide for comparing MiniMax and other models without conflating public repositories, weights, APIs, and inference servers.

11 min↗
Published Updated
Beginner

Learn the basics of training, loss and evaluation

Treat inference and training as different computations, and design evaluation before training.

8 min↗
Published Updated
Beginner

Estimate memory for 24GB and 32GB Macs

Budget weights, KV cache, working memory and the operating system separately.

11 min↗
Published Updated
Beginner

Run a small LLM locally

First observe input, output and resource use, before judging answer quality.

10 min↗
Published Updated
Intermediate → local serving evidence

Local-model runtime evidence: evaluate Strata without inheriting its claims

Read a dated local-runtime release as a configuration and evidence contract: code and weight rights, local serving, release assets, author measurements, and a bounded offline fixture.

11 min↗
Published Updated
Intermediate

Choose a GPU platform and cost model

Align runtime, storage, stopping and data-transfer assumptions before comparing prices.

12 min↗
Published Updated
Beginner

RAG: retrieve knowledge instead of packing it into weights

Separate retrieval from generation for questions that need changing facts and traceable sources.

9 min↗
Published Updated
Beginner

Choose continued pretraining and preference optimization

Identify the goal and explain why a more demanding training method is needed.

12 min↗
Published Updated
Beginner

Adapt to Japanese: vocabulary, notation and tasks

Evaluate Japanese language ability separately from domain expertise.

9 min↗
Published Updated
Intermediate

Handle medical terminology: sources, negation, units and human review

Distinguish terminology lookup and document processing prototypes from validated diagnostic use.

12 min↗
Published Updated
Beginner

Record experiments and integrate AI into a service

Turn small experiments into reproducible decisions and controlled operation.

18 min↗
Published Updated
Beginner

A dated price snapshot is not a valuation or a price history

Inspect SNDK, BE, AAOI and MU quotes with timestamps, then calculate drawdowns without inventing history.

8 min↗
Published Updated
Beginner

Memory hierarchy for AI systems: locate the stalled byte

A concrete way to distinguish registers, cache, HBM/DRAM, SSD, and networked storage before making an infrastructure claim.

8 min↗
Published Updated
Intermediate

HBM4 interfaces and packaging: more pins change the system

Understand a wider HBM4 interface, logic base die, package co-design, and the limits of multiplying pin speed.

6 min↗
Published Updated
Intermediate

NAND and SSD storage for AI: keep accelerators fed

Trace an AI data path through SSDs, host memory, decompression, and GPUs, including endurance and measurement limits.

7 min↗
Published Updated
Advanced

Invalidating a memory thesis: design a dashboard that can change your mind

A serious memory thesis names its technical and financial disconfirming observations before it becomes a narrative. HBM, NAND, packaging, and equipment each fail differently.

7 min↗
Published Updated
Beginner

Where AI clusters wait: a bottleneck map from HBM to the optical fabric

HBM, topology, oversubscription, and optical links form one queueing system. Locate the limiting resource before creating an equity narrative.

12 min↗
Published Updated
Advanced

Silicon photonics and CPO: shorten the electrical path, change the service model

Co-packaged optics moves the optical engine beside the switch ASIC. It can improve the electrical budget while changing assembly, cooling, and repair.

8 min↗
Published Updated
Intermediate

Copper, AEC, and optics: choose the reach boundary before choosing a cable

Passive copper, active electrical cable, active optical cable, and transceivers solve different channel and operations problems.

8 min↗
Published Updated
Intermediate

AAOI contracts versus revenue: trace the optical module through its accounting gates

A forecast, purchase order, shipment, acceptance and recognized revenue are separate events with different evidence.

7 min↗
Published Updated
Intermediate

Arista backend networks: translating an AI fabric into ports, links, and operations

A practical reading of a leaf-spine backend: why switch density is not deployment volume, how optics enter the topology, and how to test an operational design.

8 min↗
Published Updated
Advanced

NVLink, InfiniBand, and Ethernet: identify the domain before comparing fabrics

Interconnects operate at different layers and deployment choices. Compare topology, software, NICs, switches, and reach.

8 min↗
Published Updated
Advanced

Invalidating an optical-infrastructure thesis: make the disconfirming dashboard first

A thesis should name the technical, customer, manufacturing and accounting observations that would change it.

8 min↗
Published Updated
Beginner

Read this before adopting OSS AI: repository, weights, dependencies, and operations

Move OSS safely from learning into operations by separating licenses, maintenance, reproduction, vulnerabilities, and model weights instead of trusting star counts.

11 min↗
Published Updated
Beginner

OSS release observation: a release page is a verification starting point

A dated GitHub Releases observation with a concrete pre-adoption record and test procedure.

7 min↗
Published Updated
5 min↗
Published Updated
Intermediate

Trace one operation from end to end

Build a path of three to five arrows from the entry point to the external boundary.

5 min↗
Published Updated
Intermediate → serving design

vLLM 0.29–0.30: treat a runner default as a serving-boundary change

A dated release note on testing vLLM runner, queue, and route authorization changes before adopting a new serving release.

8 min↗
Published Updated
Intermediate → reliability evaluation

llama.cpp b11377: make JSON-schema output a parser-specific contract

A prerelease fixes a Ling 3.0 response-format gap; adopt it only through a pinned, parser-specific schema fixture.

10 min↗
Published Updated
Intermediate

Build safely and pass one test

Distinguish environment differences from problems in the code

5 min↗
Published Updated
Intermediate → decision evaluation

llama.cpp /v1/systemone: use a local decision endpoint as a routing signal

A dated local GGUF decision endpoint returns typed probabilities; it narrows routing, but never grants authority.

9 min↗
Published Updated
Intermediate

Use a debugger to test one hypothesis

Decide where to pause, which values to inspect, and what you expect

5 min↗
Published Updated
Intermediate → GPU failure diagnosis

llama.cpp CUDA MoE: treat a temporary-buffer dimension as a test boundary

A narrowly scoped prerelease fix shows why an MoE CUDA fixture must pin tensor shape, backend, and failure state before it changes a deployment.

8 min↗
Published Updated
24 min↗
Published Updated
Beginner

Read AI papers into implementation decisions: claims, conditions, reproduction, and failure

A textbook for turning papers into practical experiments by separating a claimed improvement from its problem, conditions, reproduction, trade-offs, and failures.

10 min↗
Published Updated
Intermediate

Read DeepSeek-R1: what reasoning RL changes and what it does not promise

Read the 2025 DeepSeek-R1 paper through reward design, distillation, evaluation conditions, reproduction, and operational limits.

10 min↗
Published Updated
Intermediate

Read s1: what to measure when improving reasoning with little data

Translate the 2025 s1 paper into implementation decisions through test-time scaling, budget forcing, data selection, reproduction, and limits.

9 min↗
Published Updated
Intermediate

Read DAPO: do not view LLM reinforcement learning as an algorithm alone

Read the 2025 DAPO paper through reward, length bias, asynchronous system design, reproduction, and reward hacking.

9 min↗
Published Updated
Intermediate

Read 2025 reasoning-RL and efficiency papers into implementation

Read four 2025 papers on test-time compute allocation, RL training efficiency, shorter reasoning, and dynamic length control by separating claims from constraints.

10 min↗
Published Updated
Evaluation design → offline repeat fixture

Agent evaluation: repeat the configured system before crediting the model

A bounded offline worksheet for measuring a specific agent configuration, its verifier authority, and repeat-to-repeat variation.

9 min↗
Published Updated
Intermediate

2025 VLA map: read SmolVLA, π0/π0.5, OpenVLA, and GR00T using the same yardstick

Updated notes that interpret each presentation in terms of openness, body, data, behavioral expression, and evaluation boundaries, rather than ranking each presentation in a comparison table.

12 min↗
Published Updated
Intermediate

LeRobot practice: Make dataset, imitation learning, evaluation, and Sim2Real into one quality loop

Starting with LeRobot's data format and unified interface, learn the points of contact that are likely to fail from collection to actual machine evaluation.

15 min↗
Published Updated
Beginner

Starting Physical AI without a robot: Safe learning procedure and decision not to proceed to the actual machine

It clearly states that the hardware is not running, and touches on the core of Physical AI just through data observation, simulation, and evaluation design.

15 min↗
Published Updated
Intermediate

Reading VLA papers to implementable designs: Comparative reading exercise of π0, SmolVLA, and GR00T

The novelty of a paper is not confused with performance ranking, but is broken down into behavioral expression, data, body, evaluation, and unverified boundaries.

16 min↗
Published Updated
Intermediate

Physical AI workshop: from data splits to policy evaluation and Sim2Real decisions

Turn public or synthetic episodes into reproducible splits, baseline and adaptation evaluations, failure analysis, and a justified decision not to proceed to hardware.

12 min↗
Published Updated
Evaluation fixture → offline safety boundary

Action chunks: measure recovery after a small perturbation

A strictly offline fixture for separating error propagation in a replayed action chunk from recovery after replanning.

9 min↗
Published Updated
Beginner

From chip to grid: map the power chain before choosing an AI-data-center thesis

A practical map of the electrical, thermal, contractual, and reliability boundaries between a GPU and a power plant.

8 min↗
Published Updated
Beginner

Bloom Energy solid-oxide fuel cells: mechanism, fuel boundary, and operating questions

How to evaluate an onsite solid-oxide fuel-cell system without confusing its product description with a lifecycle result.

7 min↗
Published Updated
7 min↗
Published Updated
6 min↗
Published Updated
Intermediate

Liquid cooling and rack density: thermal design is an interface problem

Explain direct-to-chip and liquid loops through heat transfer, water boundaries, controls, and serviceability.

6 min↗
Published Updated
Beginner

Choose AI providers as operating layers, not a capability table

Design OpenAI, Anthropic, Google, ElevenLabs, Databricks, Vercel, and Cloudflare as model, data, execution, voice, and delivery layers.

7 min↗
Published Updated
Intermediate

Evaluation, safety, and cost design before placing generative AI in production

Connect evidence, authority, recovery, observability, and updates before measuring model cleverness, so answer quality remains reproducible in operation.

8 min↗
Published Updated
Intermediate

Data-connected AI with Databricks: Agent Bricks and Unity Gateway in practice

Design authorization, meaning, audit, and cost before retrieval quality when connecting AI to enterprise data.

8 min↗
Published Updated
Intermediate

AI apps on Vercel: put AI SDK, Gateway, and Workflow at product boundaries

Use Vercel’s agentic-infrastructure announcements to design streaming, tool calls, durable jobs, secrets, and evaluation as one web product.

8 min↗
Published Updated
Intermediate

Voice AI with ElevenLabs: design consent, latency, and verification before voice quality

Learn consent, rights, explicit failure handling, latency, and evaluation for TTS, voice conversation, and dubbing from official ElevenLabs material and practical exercises.

8 min↗
Published Updated
Beginner

Official announcement watch: return provider updates to product decisions

A 2026-10-04 record of official blog and documentation entry points that keeps announcement dates, observed facts, implications, and unknowns separate.

8 min↗
Published Updated
Intermediate

30-second generation video workshop: Verify each cut and get closer to completion

A 30-second short film is broken down into small production experiments that can be reproduced from planning, cut design, generation, failure evaluation, and editing.

17 min↗
Published Updated
Intermediate

Video model research notes: Provision formats and production choices after 2025

Check Veo, Sora, Wan, and LTX in official documents and compare cloud provision, open weights, inference code, and safety measures on different axes.

14 min↗
Published Updated
Evaluation design → offline review fixture

Video-world evaluation: did the clip render the event the program recorded?

Use a replayable event timeline to test a generated clip separately from visual plausibility, with an offline N=1 review fixture.

10 min↗
Published Updated
Introductory → production and delivery decisions

Think about music production in six stages

Separate composition, arrangement, performance, rendering, mixing and mastering so a revision returns to the stage that can fix it.

10 min↗
Published Updated
Introductory → production and delivery decisions

Turn a music prompt into conditions you can hear

Change one condition at a time, inspect the full cue and keep timestamped acceptance reasons without inventing an A/B result.

11 min↗
Published Updated
Introductory → production and delivery decisions

Choose local generation or Suno, then inspect the export

Compare execution responsibility, editable outputs and download conditions; distinguish sample rate, codec, container and loudness.

12 min↗
Published Updated
Introductory → production and delivery decisions

Receive audio and prepare a video for publication

Track permission, source materials and download terms before preparing a YouTube delivery and testing playback on the receiving page.

13 min↗
Published Updated
Beginner

Y Combinator observation: turn public signals into constrained experiments

A dated reading of YC RFS and company material, separated from recommendations and adoption claims.

11 min↗