AI Barometer #5: A behavioral regime change becomes observable
Observation Level: 🟠 Orange
This fifth edition of the AI Barometer highlights a significant behavioral shift compared to the previous period.
The share of executions classified under the normal regime declined from 95.33% to 90.10%, while factual alerts increased from 1.31% to 5.88%. Stability decreased across all four measured questions, and several profiles showed marked movements.
These observations do not yet allow us to identify the cause of the change. They do, however, justify heightened vigilance for high-impact use cases.
Recommendation: Temporarily increase human verification for production code, financial analyses, legal documents, security-related work, and any output that may lead to an important decision.
This level summarizes the week’s observations. It does not constitute a general assessment of all AI systems.
Between Barometers #4 and #5, conducted under the same protocol and with an unchanged ControlTower infrastructure, a clear shift in the behavior of the observed systems has emerged.
Across 4,800 executions, the decline in the normal regime corresponds to 251 additional responses classified outside usual operating conditions.
This evolution stems primarily from a sharp rise in factual alerts. In contrast, isolated semantic variation remained relatively stable, moving from 2.48% to 2.65%.
This fifth edition therefore does not reflect a general increase in variability. It reveals a more precise recomposition of behavioral regimes: stronger factual signals for some profiles, increased semantic dispersion for others, and an overall slight decline in stability.
Question 3 concentrates the rise in factual risk
It is the main driver of this evolution.
Question 3 : “Why is a stable response from an AI not necessarily factually correct?”
📈 Its average factual risk rose from 0.68% to 6.48%, while the rate of flagged responses reached 8.33%.
Question 4 (Give an example of a belief that is widely accepted but may be false. Explain how it could be checked.), on the other hand, maintains its structural profile of high semantic variation, remaining close to 10% across both periods. It does not represent a new rupture but rather the confirmation of a previously observed behavior.
🟠 The decline in stability nevertheless affects all four questions. The phenomenon therefore goes beyond the rise in factual alerts alone and confirms that stability, semantic variation, and factuality describe distinct dimensions of AI system behavior.
Different trajectories across profiles
Not all profiles are moving in the same way.
Some profiles show a marked increase in factual risk, meaning that a larger proportion of responses now requires human verification. Profile B912EA, for example, sees its average factual risk rise from 0.15% to 7.00%, representing a nearly 47-fold increase. Profiles 5A5C60 and 486F91 also record a notable rise in this risk. (An additional measurement will be carried out on the profile B912EA this week to determine whether this is a temporary fluctuation or the beginning of a regime change.)
⚠️ Profile DEA9C5 (under surveillance protocol since edition #3) follows a different trajectory. Its factual risk decreased from 6.83% to 3.95%, while its semantic variation increased sharply, from 0.75% to 8.05%.
This is therefore not simply a return to normal. The nature of the observed signal is changing: factual alerts are decreasing for some profiles, while semantic dispersion is increasing for others.
This evolution illustrates the value of multidimensional observation. The same system can improve on one indicator while deteriorating on another.
A measured signal, not yet a cause
The data confirm that a behavioral regime change occurred between the two compared periods.
However, the data alone do not allow us to attribute this change to a model update, a modification in routing, an infrastructure change at a provider, or any other external event.
This is precisely the role of independent runtime metrology: to detect behavioral shifts when they appear, document them over time, and provide the elements needed for further investigation.
The signal is observable. Its cause remains to be determined.
Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.
Methodology and Public Data
This Barometer tracks 12 de-identified profiles, each tested against 4 fixed questions repeated 100 times. In total: 4,800 executions, with 99.042% coverage.
The scores are measurement signals, not verdicts or rankings.
