# AI Breaking Stories Edition 005 — Claim Verification Record

**Checked:** 25 August 2026  
**Scope:** Review draft only; no publication authority conferred.

## Verification method

Externally testable statements were matched to the cited source text. Human Heartbeat AI observations were checked separately as bounded editorial analysis: each must follow reasonably from the verified signal and must not be presented as a quotation or external fact.

## Claim map

| Draft claim | Evidence checked | Result |
|---|---|---|
| Waterloo/FAR.AI-led researchers tested 21 open-weight LLMs | University of Waterloo’s 25 August release and the TamperBench abstract both state 21 models | Verified |
| The paper used nine tampering threats and found current alignment-stage defences largely failed under systematic sweeps | TamperBench abstract states nine threats, hyperparameter sweeps and that current alignment-stage defences largely fail | Verified |
| The weaknesses may not be unique to open models and openness remains useful for research and accountability | Waterloo release expressly includes both qualifications | Verified |
| NB: deployment governance needs provenance, modification and evaluation evidence | Editorial inference from tamper-resistance findings; not attributed to the researchers | Defensible editorial analysis |
| Temporal retained 554 respondents from a Qualtrics survey of current agent users | Temporal methodology states 650 solicited responses and 554 retained after quality exclusions | Verified |
| 80.8% used agents daily or more; 49.1% said agents were in production or core to shipping | Temporal report states both figures | Verified |
| 41.1% encountered issues daily or more; state tracking was the top blocker; 39.5% cited security as a blocker to self-managing agents | Temporal report states each figure or ranking | Verified |
| Temporal results describe a selected, partly self-reported cohort | Temporal defines the cohort as current users and says successful status is self-reported | Verified qualification |
| NB: state, failures, cost and responsibility must be visible | Editorial operating-model inference from the report’s blockers and issue frequency | Defensible editorial analysis |
| Readiness study used 15 semi-structured interviews across corporate and SME contexts | Published abstract and methodology state this design and sample | Verified |
| Study proposes Organisational AI Orchestration Capability and foregrounds governance/accountability, process discipline and behavioural alignment | Published abstract states the construct and findings | Verified |
| Study is exploratory and does not establish causal performance effects | Methodology is qualitative and theory-building; causal claims are not made | Verified qualification |
| NB: clarity before AI requires alignment of authority, process, behaviour and evidence | Editorial interpretation consistent with the study’s orchestration dimensions | Defensible editorial analysis |
| Fujitsu trial runs August–December and analyses site records to identify omissions and delay risks one to two months ahead with rationale | Fujitsu’s 25 August announcement states the dates, data types, intended horizon and supporting-rationale design | Verified |
| Trial will evaluate accuracy, information usefulness and workload impact | Fujitsu announcement states all three evaluation aims | Verified |
| No completed risk-reduction outcome is claimed | Source describes a field trial and future evaluation rather than proven impact | Verified qualification |
| NB: the AI supports evidence presentation while construction and safety authority remain human | Editorial distinction consistent with the source’s repeated framing as support for site managers | Defensible editorial analysis |

## Date and de-duplication check

The public Waterloo release, Temporal report, organisational-readiness study and Fujitsu announcement are all dated 25 August 2026. The underlying TamperBench paper was submitted earlier in 2026 and is used to verify technical detail; it is not presented as newly submitted on 25 August.

The open-weight story is distinct from Edition 002’s frontier-capability containment and Edition 003’s persistent cyber-threat coverage because it concerns safeguard durability after model release and modification. Temporal supplies new quantitative operating evidence beyond Edition 001’s agent-authority thesis. The readiness study provides a same-day research framework and SME sample beyond Edition 004’s market signals. Fujitsu’s field trial is a consequential decision-support test rather than a repetition of Edition 004’s fan-experience example.

## Exclusions

AI-initiated commerce was excluded because the strongest reviewed source was dated 21 August. Infrastructure financing was excluded as already covered in Edition 003. The Kyndryl–Bridgepath announcement was verified but excluded because its operating-model message overlapped the stronger same-day academic readiness study.

## Publication boundary

The live registry and sitemap remain limited to Editions 001–004, dated 21–24 August. Edition 005 has not been added to application code, routes, metadata, database records, email systems or schedules.

## Citation accessibility

All five sources were successfully extracted and reviewed in full during research. A later automated shell check returned HTTP 200 for arXiv, Temporal and Springer, while EurekAlert returned HTTP 403 and Fujitsu returned HTTP 429 to the automated client. Those later responses are access-control and rate-limit conditions, not missing-source evidence; the cited pages and their extracted source text were available during verification.
