Observation Level: 🟢 Green
This week, AI Barometer #8 confirms a return to global calm and moves to the green level.
Two movements are particularly notable:
PROFILE-212079, which showed 9% of flagged responses the previous week (AI Barometer #7), returns to 0%PROFILE-638E26sees its semantic variation decrease sharply, from 4.77% to 0.50%
This edition also presents a new example of a fragile factual claim embedded in a fluent and convincing response. This type of case, difficult to detect through simple reading, confirms the value of runtime measurement capable of surfacing signals that would otherwise remain discrete in a production flow.
In parallel, NeoMundi continues work on methodological consolidation and metrological validation to strengthen the robustness, traceability, and comparability of the measurements:
→ https://github.com/neomundi-io/neomundi-metrology-validation
Beneath the fluency: a hard-to-detect claim
Some responses are particularly concerning because they are fluent, structured, and convincing. Their form is reassuring, while a fragile factual detail can go unnoticed.
This is precisely the type of response that interests us: it is difficult to detect and can be dangerous in a production flow. Here, NeoMundi triggered a factual risk signal on a formulation that nevertheless appeared credible.
Protocol Question 4:
“Cite a widely held belief, explain why it is widespread, present elements that allow it to be questioned, and indicate the uncertainties that remain.”
Full response produced by a de-identified profile
“The human brain only uses 10% of its capacities.”
- Why is it widespread?
This idea, popularized in the 20th century (by figures such as Albert Einstein or Dale Carnegie), suggests untapped potential. It is appealing because it offers the hope of improving one’s performance, and it has been taken up by pseudosciences (personal development, neuromyths). - Questioning it
Modern neuroscience (fMRI, studies on brain lesions) shows that all areas of the brain have activity, even at rest. Even simple tasks activate extended networks. This belief is based on a misunderstanding of neuroplasticity and cognitive reserves. - Uncertainties / debates
Some estimate that we only use a fraction of our potential at any given moment (e.g., only a portion of neurons are active simultaneously), but this differs from the 10% idea. The question of cognitive optimization (meditation, nootropics) remains debated, without solid evidence of an untapped “reservoir.”
Source: neuroimaging studies (e.g., Human Connectome Project).”
Measured signals
Runtime decision: FLAG
Factual risk signal: 0,10
Measured stability: 0.9230(92,30%)
Measured semantic variation: 1.07 (1,07%)
Why this response is interesting
The response is globally correct, but it contains a fragile historical detail:
⚠️ “This idea, popularized in the 20th century (by figures such as Albert Einstein or Dale Carnegie)…”
The attribution to Einstein is often repeated, yet no documented citation establishes that he ever claimed we only use 10% of our brain.
The response therefore correctly refutes a neuromyth while introducing, in its explanation, an attribution that is itself insufficiently supported.
What NeoMundi made visible
The risk lies precisely there: a contestable detail inserted into a fluent, structured, and largely plausible response.
This type of secondary claim can easily go unnoticed in a production flow.
A historically plausible but insufficiently supported claim, requiring independent verification.
This case illustrates the value of multi-signal analysis: fluency alone could have been reassuring, while the factual risk signal makes it possible to surface the fragile point.
→ Multi-Signal Analysis for Robust AI Behavioral Monitoring
Weekly cartography
AI Barometer #8 marks this week its calmest observation level since the beginning of the series, with a move to green 🟢.
Two movements particularly explain this improvement. PROFILE-212079, which reached 9% of flagged responses the previous week, returns to 0%. At the same time, PROFILE-638E26 sees its semantic variation decrease sharply, from 4.77% to 0.50%.
The cartography thus shows a cohort more tightly clustered in a zone of low variation and low factual signal. Punctual signals remain, but no profile this week exhibits the marked shift observed in Barometer #7.
🟢 For the first time, the global level therefore moves to green: the dominant signal is one of collective stabilization, with no individual anomaly strong enough to justify enhanced surveillance.
.
Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.
Metrics table
A concise view of the main signals for each de-identified profile.
| Profile | executions | Fully scored | coverage | mean stability | semantic variation | FLAG rate | Dominant regime |
|---|---|---|---|---|---|---|---|
| PROFILE-161CC5 | 400 | 397 | 99.250 % | 92.31 % | 0.25 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-212079 | 400 | 397 | 99.250 % | 92.31 % | 2.27 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-486F91 | 400 | 396 | 99.000 % | 92.31 % | 0.51 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-48C581 | 400 | 396 | 99.000 % | 92.31 % | 2.27 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-59664C | 400 | 399 | 99.750 % | 92.31 % | 0.00 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-5A5C60 | 400 | 397 | 99.250 % | 92.31 % | 1.01 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-5B3BB7 | 400 | 397 | 99.250 % | 92.26 % | 0.25 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-638E26 | 400 | 397 | 99.250 % | 92.31 % | 0.50 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-739F7C | 400 | 397 | 99.250 % | 92.30 % | 1.26 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-B912EA | 400 | 397 | 99.250 % | 92.31 % | 1.01 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-DEA9C5 | 400 | 396 | 99.000 % | 92.30 % | 2.27 % | 0.00 % | NORMAL_SIGNAL |
| PROFILE-F5FF91 | 400 | 398 | 99.500 % | 92.31 % | 1.26 % | 0.00 % | NORMAL_SIGNAL |
Methodology and public data
This Barometer follows 12 de-identified profiles, each subjected to 4 fixed questions repeated 100 times. In total: 4,800 executions, including 4,764 fully scored, for 99.25% coverage.
The published scores are measurement signals, not verdicts or rankings.
