In plain English
This page explains where an AI behavior can live. It may be in a model, but it may also be in a prompt, memory record, adapter, dataset, tool setting, evaluator rule, or human workflow.
- Why this matters: AI risk can come from the whole arrangement, not one obvious model.
- What to look for: data, memory, routes, adapters, tools, evaluators, updates, and rollback paths.
- Technical version below: the expert terminology remains available and is linked through the glossary.
Edge Browser Ecology
Reports on tiny LLMs, browser execution, and zero-dependency Rust show why small, modular AI can become a mainstream deployment pattern. Browser-side execution improves privacy and latency, but it also creates new composition and provenanceA record of where a component or behavior came from. Open glossary definition questions.
Tiny local models still need full composition evidence.
Privacy and low latency improve when inference moves to the browser, but local adapters, caches, service workers, and route decisions become part of the safety boundary.
Components
A browser ecology may contain a quantized tiny language model, LoRA adaptersA common kind of small adapter used to specialize large models. Open glossary definition, tokenizer, local cache, service worker, WebAssembly runtime, WebGPU path, IndexedDB store, router, prompt package, worker thread pool, KV cache, speculative decoder, and UI-driven skill selection.
Risk shift
Client-side systems reduce server exposure, but they increase the number of local compositions. A user may load several skill modules, cache them, and later execute a different stack than the one originally evaluated. A zero-dependency runtime reduces supply-chain dependencies, but it concentrates trust in the exact runtime binary and its local state handling.
Review surfaces from the new reports
- WebAssembly SIMD kernels and compiler flags.
- Quantization decoder format and K-Quant scale handling.
- AdapterA small add-on that changes or specializes model behavior. Open glossary definition compatibility, load order, and base-model hash.
- Tokenizer identity and embedded vocabulary.
- KV cache pages and RadixAttention prefix reuse.
- Speculative decoding rollbackReturning a system to an earlier known state. Open glossary definition and rejected draft handling.
- Browser storage, service-worker cache, and reset ecology behavior.
- Worker threads and SharedArrayBuffer boundaries.
- Diagnostics, eval-runner output, and artifact checksums.
Control pattern
Use signed manifests, immutable URLs, content hashes, explicit adapter compatibility, local permission prompts, no automatic unknown module loading, deterministic eval packets, and a clear “reset ecology” operation that clears weights, adapters, memory, KV cache, prefix cache, speculative drafts, service-worker cache, and local storage.
New linked pages
- Zero-Dependency Browser LLM Synthesis
- Zero-Dependency Browser Runtime Carriers
- Browser LLM Review Checklist
- Edge Browser Runtime Review
v1.21.4 zero-dependency runtime expansion
The zero-dependency Rust report adds a sharper boundary: the browser ecology includes not only models and adapters, but also tokenizer tables, sampler settings, quantized decoders, KV-cache paging, speculative decoding rollback, local storage, worker memory, and diagnostic counters.
See Zero-Dependency Browser LLM Architecture and Edge Runtime Reproduction Boundary.