Four developments reported on 25 August sharpen the same operating question. A model can begin with safeguards and still be modified. An agent can feel productive while its state and failures remain difficult to track. An organisation can own the technology while lacking the governance and behavioural alignment to use it well. A decision-support system can surface useful risks while leaving the real judgement exactly where it belongs: with a competent human.
Safeguards have to survive contact with the deployed model
The University of Waterloo reported that an international team led by Waterloo and FAR.AI tested 21 open-weight large language models and found that each could be modified despite its built-in protections. The underlying TamperBench paper evaluates those models across nine tampering threats using standardised safety and capability measures; its authors report that current alignment-stage defences largely failed under systematic attack sweeps.
That finding should not be distorted into a claim that openness itself is irresponsible. The researchers explicitly say that open-weight models remain important for research and accountability, and that the weaknesses may not be unique to open systems. The sharper lesson is that a protection visible in the original release cannot be assumed to survive later fine-tuning, modification or deployment.
Read source: University of Waterloo via EurekAlert — Major security weaknesses found in leading open AI models ↗Agent adoption is moving faster than operational assurance
Temporal’s 25 August State of Development report describes a survey of 554 engineers and engineering leaders who were already using AI agents. It reports that 80.8% used agents daily or more, while 49.1% said agents were in production or core to how they shipped. Yet 41.1% said they encountered agent issues daily or more, state tracking was the leading blocker to wider use, and 39.5% cited security concerns as an obstacle to self-managing agents.
These figures describe a selected, vendor-commissioned cohort rather than the whole engineering profession. ‘Successful’ adoption and many performance measures are self-reported. Even within those limits, the contrast matters: high usage does not remove the need to track state, debug failures, control costs or understand who is responsible when an agent’s work crosses into a live system.
Read source: Temporal — State of Development Report 2026 ↗AI readiness is becoming an orchestration capability
An open-access study published on 25 August argues that organisational AI readiness is shifting from a static inventory of technology and resources towards an ability to coordinate governance, processes and behaviour. Drawing on prior literature and 15 semi-structured interviews across corporate and SME contexts, the authors propose ‘Organisational AI Orchestration Capability’ as the capacity to integrate, govern and operationalise heterogeneous AI resources into reliable outcomes.
The study is exploratory. Its small qualitative sample does not prove that the proposed framework causes better performance. Its practical contribution is the distinction it makes: as infrastructure and data become baseline prerequisites, governance and accountability, process discipline and behavioural alignment become the harder differentiators.
Read source: Discover Artificial Intelligence — The evolution of organisational AI readiness toward an orchestration capability ↗Decision support becomes credible when it shows its working
Fujitsu, Tokyu Construction and Kitano Construction announced a field trial at Fujitsu Technology Park that analyses schedules, daily reports, procedures, inspections and related site data. The system is intended to identify possible omissions and delay risks one to two months ahead and present supporting rationale to site managers. The trial, running from August to December, will evaluate detection accuracy, the usefulness of the underlying information and whether the system reduces managers’ workload.
This is not evidence that the system has already prevented delays or safety failures. It is a test of whether evidence-grounded AI can help experienced people see a developing problem earlier across complex, interdependent work. The responsible design signal is that the system surfaces a risk and its basis; the site manager remains responsible for the judgement and action that follow.
Read source: Fujitsu — AI-supported construction process and risk management field trial ↗