AnatomyStrong architectural inferencev1.22.1

In plain English

This page explains where an AI behavior can live. It may be in a model, but it may also be in a prompt, memory record, adapter, dataset, tool setting, evaluator rule, or human workflow.

  • Why this matters: AI risk can come from the whole arrangement, not one obvious model.
  • What to look for: data, memory, routes, adapters, tools, evaluators, updates, and rollback paths.
  • Technical version below: the expert terminology remains available and is linked through the glossary.

Edge Browser Ecology

Evidence levelStrong architectural inferenceTechnical label: Architectural inference

Reports on tiny LLMs, browser execution, and zero-dependency Rust show why small, modular AI can become a mainstream deployment pattern. Browser-side execution improves privacy and latency, but it also creates new composition and A record of where a component or behavior came from. Open glossary definition questions.

schematic · edge/browser model ecology

Tiny local models still need full composition evidence.

Privacy and low latency improve when inference moves to the browser, but local adapters, caches, service workers, and route decisions become part of the safety boundary.

Components

A browser ecology may contain a quantized tiny language model, A common kind of small adapter used to specialize large models. Open glossary definition, tokenizer, local cache, service worker, WebAssembly runtime, WebGPU path, IndexedDB store, router, prompt package, worker thread pool, KV cache, speculative decoder, and UI-driven skill selection.

Risk shift

Client-side systems reduce server exposure, but they increase the number of local compositions. A user may load several skill modules, cache them, and later execute a different stack than the one originally evaluated. A zero-dependency runtime reduces supply-chain dependencies, but it concentrates trust in the exact runtime binary and its local state handling.

Review surfaces from the new reports

Control pattern

Use signed manifests, immutable URLs, content hashes, explicit adapter compatibility, local permission prompts, no automatic unknown module loading, deterministic eval packets, and a clear “reset ecology” operation that clears weights, adapters, memory, KV cache, prefix cache, speculative drafts, service-worker cache, and local storage.

New linked pages

v1.21.4 zero-dependency runtime expansion

Evidence levelStrong architectural inferenceTechnical label: Architectural inference

The zero-dependency Rust report adds a sharper boundary: the browser ecology includes not only models and adapters, but also tokenizer tables, sampler settings, quantized decoders, KV-cache paging, speculative decoding rollback, local storage, worker memory, and diagnostic counters.

See Zero-Dependency Browser LLM Architecture and Edge Runtime Reproduction Boundary.