CompositionStrong architectural inferencev1.22.1

In plain English

This page explains why testing AI parts one by one is necessary but incomplete. Safe-looking parts can still produce unsafe behavior when combined.

  • Why this matters: AI risk can come from the whole arrangement, not one obvious model.
  • What to look for: data, memory, routes, adapters, tools, evaluators, updates, and rollback paths.
  • Technical version below: the expert terminology remains available and is linked through the glossary.

The Unsafe Unit Is the Composition

Evidence levelStrong architectural inferenceTechnical label: Architectural inference

The unsafe unit is sometimes not the model, adapter, prompt, or A system that judges whether an AI output or candidate is acceptable. Open glossary definition alone. It is the relationship among them under a specific route and state.

Mechanism

The mechanism is interaction. Components exchange context through hidden state, prompts, outputs, adapters, memory retrieval, tool calls, evaluator prompts, and release rules. Each interaction can change what the next component sees and what the system is allowed to do.

Evaluation implication

The evidence record should include the exact A machine-readable record of the exact runtime composition used for an evaluation, release, incident, or rollback. Open glossary definition. A statement such as “Adapter C passed” is incomplete unless it says which base model, load order, router, prompt package, memory snapshot, evaluator, inference configuration, and deployment environment were used.

Practical control

Use composition-aware test suites, targeted higher-order samples, route-level canaries, independent judges, and Returning a system to an earlier known state. Open glossary definition packets that include all relevant runtime dependencies.

<!-- expanded-release-content -->

What counts as the unit

Evidence levelStrong architectural inferenceTechnical label: Architectural inference

A composition includes the base model, adapters, merge coefficients, load order, prompt package, router policy, memory state, tool profile, The exact version of the evaluator used for a test or release. Open glossary definition, inference settings, deployment environment, and release alias. It also includes time: which component wrote a memory, which evaluator approved it, and which descendant later consumed it.

A model-only unit is sometimes adequate. It is inadequate when behavior appears only after components interact. In those cases, the system can pass every isolated component check while failing at the composed boundary. The failure is not mysterious; the tested object and the deployed object were different.

Why averages hide it

Average benchmark scores can improve while a narrow composition-specific failure appears. A routing system may send only a small fraction of tasks through the risky path. A memory trigger may appear only after prior conversations. A tool permission may matter only for one class of users. A judge may evaluate the final answer without seeing the hidden decomposition that produced it.

Documentation requirement

Provenance should include runtime composition, not only artifact lineage. A A visual or machine-readable map of derivation history. Open glossary definition says where components came from. A composition manifest says what was actually loaded, in what order, with what permissions, under what evaluator, at what UTC time. Both are needed.

Review prompt

When a release is approved, ask whether the approval applies to the exact composition being deployed. If the answer is “approximately,” the residual risk should be recorded rather than hidden under the model name.