REGISTERED TAG
Evaluation
Define conditions and criteria; distinguish estimates from measured outcomes.
Version agent work with Cloudflare Artifacts
Separate Git history, execution state, and large outputs; scope credentials and validate the exact commit before accepting agent work.
Design AI agents that survive disconnects
Separate a browser connection, durable work, and stored conversation state. Use PiHarness to reason about admission, reconnects, replay, and permission boundaries.
Access service authentication: verify the machine principal on every request
Use an offline request matrix to separate a service credential, a browser session, an edge decision, and an authorized application action.
Computer use: build systems that can observe, act, and recover
A practical map of browser and desktop automation: perception, planning, action, verification, and safe recovery.
Understand Skills by separating them from other mechanisms
Learn what changes and what stays the same, rather than memorizing names.
Antigravity 09-2026: file edits, hook coverage, and stale-state rejection
A new file-tool contract changes both dispatch and authorization; rehearse one edit and reject stale state before migrating.
Separate discovery, reading, and use
Understanding loading stages helps you isolate why a Skill is not working.
Read safety, licenses, and dependencies
A Skill is text to read and can also lead to code and external connections.
Create a small Skill: review a lesson against its sources
Reduce one failure of your own before adding many ready-made Skills.
Compare with and without a Skill fairly
Explain what your experiment changed instead of relying on one impression.
Read differences and retain them in an experiment ledger
Move from feeling that things improved to explaining what changed.
Maintain comparisons across updates
Keep Skills in maintainable units rather than simply adding more.
What will Skills do as models improve?
Turn possible future value into testable questions instead of predictions stated as facts.
Databricks AI roles: Genie, governed agents, memory and evaluation
Assign business questions, agent workflows and coding assistance to the right roles. Follow tool identities, state isolation, budget limits and failure evidence in a support case.
Hacker News observation: turn attention into bounded technical experiments
A dated HN API snapshot, read as a set of technical hypotheses rather than a popularity ranking.
A map of creative deliverables and work
Distinguish a single image from an interactive service.
Use AI to propose candidates under constraints
Separate the model, inputs, generation, editing and verification.
Turn visual direction into words and specifications
Give instructions through composition, color, light, shape and purpose.
Practical image workshop: Creating a single advertising visual using a verifiable process
An exercise in creating still images containing products and people by dividing them into specifications, references, composition, creation, local corrections, typesetting, and release checks.
Produce motion and video through separate workflows
Work with timing, states, editing and sound
Measure quality and cost to guide the next step
Keep tokens, images, seconds and credits distinct
Jev foundations: typed decisions outside text generation
Use Jev/System One as a small typed decision component with candidates, state, and probabilities alongside ordinary code.
Jev Choice, Score, and Noul: do not confuse probability, confidence, and authority
Learn the roles of three question primitives and how to prevent concentrated distributions from becoming permission to automate.
Jev implementation lab: connect candidate generation, abstention, and E2E evaluation
Connect semantic judgment to deterministic candidate generation and safety boundaries, then design end-to-end evaluation with controls.
Map AI, machine learning and LLMs
Match the problem you want to solve to a mechanism, rather than treating AI as one uniform tool.
Quantization, KV cache, and runtimes for a local LLM that stays responsive
Separate model weights, working memory, and generation speed; choose MLX, llama.cpp, or vLLM and turn the choice into a reproducible local evaluation.
Understand models, tokens and inference
Follow the transformation from text to numbers, then from numbers to the next token.
Choosing MiniMax and local LLMs: separate open weights, APIs, and feasibility
A practical research and experiment guide for comparing MiniMax and other models without conflating public repositories, weights, APIs, and inference servers.
Learn the basics of training, loss and evaluation
Treat inference and training as different computations, and design evaluation before training.
Estimate memory for 24GB and 32GB Macs
Budget weights, KV cache, working memory and the operating system separately.
Run a small LLM locally
First observe input, output and resource use, before judging answer quality.
Local-model runtime evidence: evaluate Strata without inheriting its claims
Read a dated local-runtime release as a configuration and evidence contract: code and weight rights, local serving, release assets, author measurements, and a bounded offline fixture.
Choose a GPU platform and cost model
Align runtime, storage, stopping and data-transfer assumptions before comparing prices.
RAG: retrieve knowledge instead of packing it into weights
Separate retrieval from generation for questions that need changing facts and traceable sources.
Choose continued pretraining and preference optimization
Identify the goal and explain why a more demanding training method is needed.
Adapt to Japanese: vocabulary, notation and tasks
Evaluate Japanese language ability separately from domain expertise.
Handle medical terminology: sources, negation, units and human review
Distinguish terminology lookup and document processing prototypes from validated diagnostic use.
Record experiments and integrate AI into a service
Turn small experiments into reproducible decisions and controlled operation.
A dated price snapshot is not a valuation or a price history
Inspect SNDK, BE, AAOI and MU quotes with timestamps, then calculate drawdowns without inventing history.
Memory hierarchy for AI systems: locate the stalled byte
A concrete way to distinguish registers, cache, HBM/DRAM, SSD, and networked storage before making an infrastructure claim.
HBM4 interfaces and packaging: more pins change the system
Understand a wider HBM4 interface, logic base die, package co-design, and the limits of multiplying pin speed.
NAND and SSD storage for AI: keep accelerators fed
Trace an AI data path through SSDs, host memory, decompression, and GPUs, including endurance and measurement limits.
Invalidating a memory thesis: design a dashboard that can change your mind
A serious memory thesis names its technical and financial disconfirming observations before it becomes a narrative. HBM, NAND, packaging, and equipment each fail differently.
Where AI clusters wait: a bottleneck map from HBM to the optical fabric
HBM, topology, oversubscription, and optical links form one queueing system. Locate the limiting resource before creating an equity narrative.
Silicon photonics and CPO: shorten the electrical path, change the service model
Co-packaged optics moves the optical engine beside the switch ASIC. It can improve the electrical budget while changing assembly, cooling, and repair.
Copper, AEC, and optics: choose the reach boundary before choosing a cable
Passive copper, active electrical cable, active optical cable, and transceivers solve different channel and operations problems.
AAOI contracts versus revenue: trace the optical module through its accounting gates
A forecast, purchase order, shipment, acceptance and recognized revenue are separate events with different evidence.
Arista backend networks: translating an AI fabric into ports, links, and operations
A practical reading of a leaf-spine backend: why switch density is not deployment volume, how optics enter the topology, and how to test an operational design.
NVLink, InfiniBand, and Ethernet: identify the domain before comparing fabrics
Interconnects operate at different layers and deployment choices. Compare topology, software, NICs, switches, and reach.
Invalidating an optical-infrastructure thesis: make the disconfirming dashboard first
A thesis should name the technical, customer, manufacturing and accounting observations that would change it.
Read this before adopting OSS AI: repository, weights, dependencies, and operations
Move OSS safely from learning into operations by separating licenses, maintenance, reproduction, vulnerabilities, and model weights instead of trusting star counts.
OSS release observation: a release page is a verification starting point
A dated GitHub Releases observation with a concrete pre-adoption record and test procedure.
Turn a moving branch into a fixed reference
Make reading notes with commits and sources attached.
Trace one operation from end to end
Build a path of three to five arrows from the entry point to the external boundary.
vLLM 0.29–0.30: treat a runner default as a serving-boundary change
A dated release note on testing vLLM runner, queue, and route authorization changes before adopting a new serving release.
llama.cpp b11377: make JSON-schema output a parser-specific contract
A prerelease fixes a Ling 3.0 response-format gap; adopt it only through a pinned, parser-specific schema fixture.
Build safely and pass one test
Distinguish environment differences from problems in the code
llama.cpp /v1/systemone: use a local decision endpoint as a routing signal
A dated local GGUF decision endpoint returns typed probabilities; it narrows routing, but never grants authority.
Use a debugger to test one hypothesis
Decide where to pause, which values to inspect, and what you expect
llama.cpp CUDA MoE: treat a temporary-buffer dimension as a test boundary
A narrowly scoped prerelease fix shows why an MoE CUDA fixture must pin tensor shape, backend, and failure state before it changes a deployment.
Move on when you can explain what you learned
Turn AI guidance into your own understanding supported by evidence
Read AI papers into implementation decisions: claims, conditions, reproduction, and failure
A textbook for turning papers into practical experiments by separating a claimed improvement from its problem, conditions, reproduction, trade-offs, and failures.
Read DeepSeek-R1: what reasoning RL changes and what it does not promise
Read the 2025 DeepSeek-R1 paper through reward design, distillation, evaluation conditions, reproduction, and operational limits.
Read s1: what to measure when improving reasoning with little data
Translate the 2025 s1 paper into implementation decisions through test-time scaling, budget forcing, data selection, reproduction, and limits.
Read DAPO: do not view LLM reinforcement learning as an algorithm alone
Read the 2025 DAPO paper through reward, length bias, asynchronous system design, reproduction, and reward hacking.
Read 2025 reasoning-RL and efficiency papers into implementation
Read four 2025 papers on test-time compute allocation, RL training efficiency, shorter reasoning, and dynamic length control by separating claims from constraints.
Agent evaluation: repeat the configured system before crediting the model
A bounded offline worksheet for measuring a specific agent configuration, its verifier authority, and repeat-to-repeat variation.
2025 VLA map: read SmolVLA, π0/π0.5, OpenVLA, and GR00T using the same yardstick
Updated notes that interpret each presentation in terms of openness, body, data, behavioral expression, and evaluation boundaries, rather than ranking each presentation in a comparison table.
LeRobot practice: Make dataset, imitation learning, evaluation, and Sim2Real into one quality loop
Starting with LeRobot's data format and unified interface, learn the points of contact that are likely to fail from collection to actual machine evaluation.
Starting Physical AI without a robot: Safe learning procedure and decision not to proceed to the actual machine
It clearly states that the hardware is not running, and touches on the core of Physical AI just through data observation, simulation, and evaluation design.
Reading VLA papers to implementable designs: Comparative reading exercise of π0, SmolVLA, and GR00T
The novelty of a paper is not confused with performance ranking, but is broken down into behavioral expression, data, body, evaluation, and unverified boundaries.
Physical AI workshop: from data splits to policy evaluation and Sim2Real decisions
Turn public or synthetic episodes into reproducible splits, baseline and adaptation evaluations, failure analysis, and a justified decision not to proceed to hardware.
Action chunks: measure recovery after a small perturbation
A strictly offline fixture for separating error propagation in a replayed action chunk from recovery after replanning.
From chip to grid: map the power chain before choosing an AI-data-center thesis
A practical map of the electrical, thermal, contractual, and reliability boundaries between a GPU and a power plant.
Bloom Energy solid-oxide fuel cells: mechanism, fuel boundary, and operating questions
How to evaluate an onsite solid-oxide fuel-cell system without confusing its product description with a lifecycle result.
Bloom Energy backlog and contracts: backlog is a definition with cancellation and execution risk
How to keep commercial pipeline, binding contracts, revenue, cash, and installed capacity separate.
Gas supply, emissions, and permits: onsite power has external boundaries
A site checklist for fuel contracts, air permits, methane boundaries, and emissions claims.
Liquid cooling and rack density: thermal design is an interface problem
Explain direct-to-chip and liquid loops through heat transfer, water boundaries, controls, and serviceability.
Choose AI providers as operating layers, not a capability table
Design OpenAI, Anthropic, Google, ElevenLabs, Databricks, Vercel, and Cloudflare as model, data, execution, voice, and delivery layers.
Evaluation, safety, and cost design before placing generative AI in production
Connect evidence, authority, recovery, observability, and updates before measuring model cleverness, so answer quality remains reproducible in operation.
Data-connected AI with Databricks: Agent Bricks and Unity Gateway in practice
Design authorization, meaning, audit, and cost before retrieval quality when connecting AI to enterprise data.
AI apps on Vercel: put AI SDK, Gateway, and Workflow at product boundaries
Use Vercel’s agentic-infrastructure announcements to design streaming, tool calls, durable jobs, secrets, and evaluation as one web product.
Voice AI with ElevenLabs: design consent, latency, and verification before voice quality
Learn consent, rights, explicit failure handling, latency, and evaluation for TTS, voice conversation, and dubbing from official ElevenLabs material and practical exercises.
Official announcement watch: return provider updates to product decisions
A 2026-10-04 record of official blog and documentation entry points that keeps announcement dates, observed facts, implications, and unknowns separate.
30-second generation video workshop: Verify each cut and get closer to completion
A 30-second short film is broken down into small production experiments that can be reproduced from planning, cut design, generation, failure evaluation, and editing.
Video model research notes: Provision formats and production choices after 2025
Check Veo, Sora, Wan, and LTX in official documents and compare cloud provision, open weights, inference code, and safety measures on different axes.
Video-world evaluation: did the clip render the event the program recorded?
Use a replayable event timeline to test a generated clip separately from visual plausibility, with an offline N=1 review fixture.
Think about music production in six stages
Separate composition, arrangement, performance, rendering, mixing and mastering so a revision returns to the stage that can fix it.
Turn a music prompt into conditions you can hear
Change one condition at a time, inspect the full cue and keep timestamped acceptance reasons without inventing an A/B result.
Choose local generation or Suno, then inspect the export
Compare execution responsibility, editable outputs and download conditions; distinguish sample rate, codec, container and loudness.
Receive audio and prepare a video for publication
Track permission, source materials and download terms before preparing a YouTube delivery and testing playback on the receiving page.
Y Combinator observation: turn public signals into constrained experiments
A dated reading of YC RFS and company material, separated from recommendations and adoption claims.