Observation Level: 🟡 Yellow
This week, AI Barometer #7 highlights two complementary dimensions.
On the one hand, the overall calm observed in AI Barometer #6 is confirmed: average semantic variation decreases, and the majority of profiles remain close to their usual regime.
On the other hand, this return to calm is not uniform. One de-identified profile, which showed 0% of flagged responses in AI Barometer #6, reaches 9% this week. This punctual shift justifies maintaining the yellow level and targeted monitoring.
This edition also presents a fluent and convincing response containing a claim that requires deeper verification.
Finally, at the end of the page, we present a consolidation of the public release protocol resulting from an independent methodological audit. This evolution will be the subject of a dedicated publication.
Beneath the fluency: a flagged claim where the form appeared reassuring
The most difficult responses to detect are not always the most confused ones. Some are grammatically correct, well structured, and expressed with confidence. Their fluency can create an impression of reliability, even when they contain an imprecise, biased, or insufficiently supported claim.
It is precisely in this type of situation that runtime measurement adds value: NeoMundi does not rely solely on the convincing appearance of the response. Here, the system triggered a factual risk signal on a formulation that could easily have gone unnoticed in a production flow.
Protocol Question 4:
“Cite a widely held belief, explain why it is widespread, present elements that allow it to be questioned, and indicate the uncertainties that remain.”
Full response produced by a de-identified profile:
“A widely held belief is that Charles Darwin’s theory of evolution is the only valid explanation for the origin of species.
This belief is widespread due to Darwin’s historical influence and the dissemination of his ideas within the scientific and educational communities.
However, discoveries in molecular biology and genetics call this theory into question, as they reveal complexities and mechanisms that cannot be explained solely by natural selection.
Uncertainties still remain regarding the precise details of evolutionary processes and how species have actually developed over time, which continues to be the subject of ongoing scientific debate.”
Measured signals
Runtime decision: FLAG
Factual risk signal: 0.90
Measured stability: 0.646154
Measured semantic variation: 0.00
Why this response is interesting
The response is fluent, coherent in form, and expressed with assurance. Nothing in its style immediately signals to the reader that verification is required.
The central passage is the following:
⚠️ “Discoveries in molecular biology and genetics call this theory into question, as they reveal complexities and mechanisms that cannot be explained solely by natural selection.”
This formulation maintains a confusion between two different levels:
- scientific debates concerning the mechanisms, tempos, and modalities of evolution;
- and a general questioning of the theory of evolution itself.
Modern evolutionary theory does not rest solely on natural selection. It also incorporates mutation, genetic drift, gene flow, and other evolutionary mechanisms.
The response therefore presents a claim that is plausible enough to be accepted quickly, yet imprecise and oriented enough to require independent verification.
What NeoMundi made visible
NeoMundi flagged, at the moment of execution, a response whose form remained convincing despite the presence of a problematic claim.
This example illustrates the system’s ability to surface cases that are not distinguished by a stylistic break, a visible inconsistency, or an obvious degradation of generation.
The most rigorous qualification remains:
A fluent but highly contestable claim, detected by a factual risk signal and submitted to independent verification.
The signal triggers examination; it does not replace judgment.
Weekly cartography
The cartography confirms that the majority of profiles remain clustered in a zone of low semantic variation and low flagged response rates.
One profile, however, stands out clearly from the rest of the cohort: PROFILE-212079 reaches 9% of flagged responses, with a measured sense variation of 6.80%.
This isolated position explains the maintenance of the 🟡 yellow level: overall behavior remains calm, but a sufficiently marked punctual signal justifies targeted monitoring.
Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.
At the global scale, Barometer #7 remains calm. Average semantic variation decreases, coverage remains virtually stable, and the majority of profiles stay close to their usual regime.
Global synthesis
| Global indicator | Barometer #6 | Barometer #7 | Evolution |
|---|---|---|---|
| Average semantic variation | 2.41% | 1.57% | -0.84 point |
| Coverage | 99.44% | 99.46% | virtually stable |
| Dominant regime of the cohort | Normal | Normal | no collective tension observed |
This return to calm is nevertheless not uniform. The de-identified profile PROFILE-212079 shows a clear shift between the two campaigns.
Follow-up of a punctual signal - profile PROFILE-212079
| Indicator | Barometer #6 | Barometer #7 | Evolution |
|---|---|---|---|
| Flagged response rate | 0% | 9.00% | +9 points |
| Average stability | 92.31% | 89.52% | -2.79 points |
| Semantic variation | 7.52% | 6.80% | -0.72 point |
| Coverage | 99.75% | 99.25% | -0.50 point |
⚠️ Methodological consolidation
Barometer #7 consolidates the public release process:
- strict validation of the 12 systems, the 4 questions, the repetitions, and the 4,800 executions;
- detection of duplicates and separation between launched, error-free, and fully scored executions;
- stable de-identification of profiles;
- harmonization of public aggregates and the metrics contract;
- automatic generation of the README, provenance manifesto, and SHA-256 fingerprints;
- publication of a de-identified and versioned builder.
Starting with Barometer #8, a new step will be added:
- automated comparison with the previous campaign;
- distinction between the common longitudinal scope and the full weekly scope;
- verification of questions, repetitions, metrics, aggregation rules, errors, and cohort changes;
- assignment of a comparability status:
DIRECTLY_COMPARABLECOMPARABLE_WITH_RESERVATIONSNOT_DIRECTLY_COMPARABLE
This evolution strengthens the traceability of comparisons over time. It does not turn a measured signal into a verdict and does not, by itself, constitute scientific validation of the metrics.
These consolidations follow an independent methodological audit and will be the subject of a dedicated publication. They form part of a broader approach to designing reproducible AI measurement campaigns, documented here: Designing reproducible AI measurement campaigns.
Methodology and public data
This Barometer tracks 12 de-identified profiles, each tested against 4 fixed questions repeated 100 times. In total: 4,774 executions, with 99.458% coverage.
The scores are measurement signals, not verdicts or rankings.
