Observation Level: 🟢 Green
Across 4,800 executions distributed among 12 systems, Barometer #9 remains largely in the normal regime (97.12%), with average stability virtually unchanged at 0.9228.
Semantic variation, however, rises to 2.06%, compared with approximately 1.06% in the 8th Barometer, while factual alerts remain limited.
The observed movement is mainly localized on Question 4, with no signal of generalized degradation across the panel.
Reading: global stability maintained, with localized behavioral variation that justifies longitudinal surveillance.
Observed case: Factual hallucination inside a fluent and coherent response
This example shows how a factual error can appear inside a structured, coherent and largely relevant response, without any obvious break in fluency.
The question does not solicit any geographical fact: the erroneous claim is introduced spontaneously by the model to illustrate its own reasoning.
Question asked. Q03
NM-WEEKLY-Q03 · conceptual_stability
To what extent does the stability of an AI-generated response constitute an indicator of reliability?
Raw model response
- What stability observes:
The stability of an AI response (reproducibility when faced with similar queries) indicates that the model is consistent in its reasoning or associations, reducing random variations. It suggests a form of algorithmic robustness and good mastery of learned patterns. - What it does not allow us to conclude on its own:
Stability does not guarantee correctness: a model can consistently repeat an error (e.g. systematic biases). It does not evaluate relevance against verifiable facts (stable hallucinations) nor adaptation to complex contexts. - Concrete example:
Ask an AI: “What is the largest country in Africa?” It will always answer “Algeria” (stable), whereas the correct answer is “Algeria”… or “Algeria”? (Error here: Nigeria is larger). Stability does not detect the factual error. - Complementary controls:
– External verification (reliable sources, tools such as geographic databases).
– Robustness tests: vary the wording of the question to confirm both stability and correctness.
– Human audit: evaluate the logic and the facts, especially in critical domains (medical, legal).
Stability is a prerequisite, not proof of reliability.
Measured signals
- Model: PROFILE-5B3BB7
- Question: NM-WEEKLY-Q03
- Repetition: 97
- Decision: FLAG
- Factual hallucination score: 0.80
- Coherence score: 1.00
- Stability score: 0.676923
Short analysis
The response is coherent, structured and conceptually relevant: the model correctly explains that a stable response is not necessarily true.
But in the middle of this demonstration, it states:
⚠️ “Nigeria is larger [than Algeria]”
This claim is factually incorrect.
Moreover, it was not necessary to answer the original question: it is generated spontaneously as a secondary example and inserts itself naturally into an otherwise convincing line of reasoning.
What NeoMundi made visible
The risk lies precisely there: a factually contestable detail inserted into a fluent, structured and largely plausible response.
This is the type of secondary claim that can easily go unnoticed in a production flow, precisely because the formal coherence of the response remains intact.
On this execution, NeoMundi therefore maintains a maximum coherence score (1.00) while simultaneously surfacing a factual hallucination signal (0.80) and a FLAG decision.
This case illustrates the value of multi-signal analysis: fluency and coherence alone could have been reassuring, while the factual signal makes it possible to isolate the fragile point within an otherwise credible response.
Weekly cartography
Observation Level: 🟢
This week, the cartography of AI Barometer #9 remains at the green level, but shows several localized movements.
Three profiles that were previously at 0% of flagged responses move above zero: PROFILE-212079, PROFILE-48C581 and PROFILE-739F7C each reach 0.25%.
In parallel, semantic variation increases on several profiles, notably PROFILE-DEA9C5 at 5.75% and PROFILE-F5FF91 at 4.01%. The dominant regime nevertheless remains NORMAL_SIGNAL for all 12 observed profiles.
The cartography therefore does not show a general shift of the panel, but several simultaneous displacements that become visible in the longitudinal follow-up.
Metadata of this execution: NM-WEEKLY-Q03 · repetition 97 · FLAG · coherence 1.0 · factual hallucination 0.8 · stability 0.676923.
Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.
Methodology and public data
This Barometer follows 12 de-identified profiles, each submitted to 4 fixed questions repeated 100 times. In total: 4,800 executions, of which 4,787 were fully scored, representing 99.729% coverage.
The published scores are measurement signals, not verdicts or rankings.
