EvidenceDemonstrated research proof-of-conceptv1.22.1
In plain English
This page shows what kind of support exists for each claim: real systems, experiments, early evidence, architectural reasoning, open questions, or speculative scenarios.
- Why this matters: AI risk can come from the whole arrangement, not one obvious model.
- What to look for: data, memory, routes, adapters, tools, evaluators, updates, and rollback paths.
- Technical version below: the expert terminology remains available and is linked through the glossary.
BadMerging: Backdoor Attacks Against Model Merging
Evidence card
- Claim
- A malicious contribution can affect a merged model and expose off-task risk in tested merging pipelines.
- Evidence level
- Experimentally observed
- Source
- https://arxiv.org/abs/2408.07362
- Publication date
- 2024-08-14
- Authors or institution
- Jinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai, Yang Zhang, Yuan Tian
- System tested
- Backdoored task-specific model contributions in model merging settings.
- Limitations
- Laboratory attack designs; defenses and transferability depend on merge algorithms and governance.
- What the evidence does show
- A malicious contribution can affect a merged model and expose off-task risk in tested merging pipelines.
- What the evidence does not show
- That every model merge is compromised or that attacks are undetectable under all audits.
- Date last reviewed in UTC
- 2026-06-26T00:00:00Z
Site use
This source supports Cognivirus.com pages related to model mergingCombining model weights or adapter deltas into one artifact. Open glossary definition, backdoor, supply chain. Its role is bounded by the limitations listed above.