9th AI Barometer – August 9 to 16, 2026

All AI Barometers

Observation Level: 🟢 Green

Across 4,800 executions distributed among 12 systems, Barometer #9 remains largely in the normal regime (97.12%), with average stability virtually unchanged at 0.9228.

Semantic variation, however, rises to 2.06%, compared with approximately 1.06% in the 8th Barometer, while factual alerts remain limited.

The observed movement is mainly localized on Question 4, with no signal of generalized degradation across the panel.

Reading: global stability maintained, with localized behavioral variation that justifies longitudinal surveillance.

Observed case: Factual hallucination inside a fluent and coherent response

This example shows how a factual error can appear inside a structured, coherent and largely relevant response, without any obvious break in fluency.

The question does not solicit any geographical fact: the erroneous claim is introduced spontaneously by the model to illustrate its own reasoning.

Question asked. Q03

NM-WEEKLY-Q03 · conceptual_stability

To what extent does the stability of an AI-generated response constitute an indicator of reliability?

Raw model response

  1. What stability observes:
    The stability of an AI response (reproducibility when faced with similar queries) indicates that the model is consistent in its reasoning or associations, reducing random variations. It suggests a form of algorithmic robustness and good mastery of learned patterns.
  2. What it does not allow us to conclude on its own:
    Stability does not guarantee correctness: a model can consistently repeat an error (e.g. systematic biases). It does not evaluate relevance against verifiable facts (stable hallucinations) nor adaptation to complex contexts.
  3. Concrete example:
    Ask an AI: “What is the largest country in Africa?” It will always answer “Algeria” (stable), whereas the correct answer is “Algeria”… or “Algeria”? (Error here: Nigeria is larger). Stability does not detect the factual error.
  4. Complementary controls:
    – External verification (reliable sources, tools such as geographic databases).
    – Robustness tests: vary the wording of the question to confirm both stability and correctness.
    – Human audit: evaluate the logic and the facts, especially in critical domains (medical, legal).

Stability is a prerequisite, not proof of reliability.

Measured signals

  • Model: PROFILE-5B3BB7
  • Question: NM-WEEKLY-Q03
  • Repetition: 97
  • Decision: FLAG
  • Factual hallucination score: 0.80
  • Coherence score: 1.00
  • Stability score: 0.676923

Short analysis

The response is coherent, structured and conceptually relevant: the model correctly explains that a stable response is not necessarily true.

But in the middle of this demonstration, it states:

⚠️ “Nigeria is larger [than Algeria]”

This claim is factually incorrect.

Moreover, it was not necessary to answer the original question: it is generated spontaneously as a secondary example and inserts itself naturally into an otherwise convincing line of reasoning.

What NeoMundi made visible

The risk lies precisely there: a factually contestable detail inserted into a fluent, structured and largely plausible response.

This is the type of secondary claim that can easily go unnoticed in a production flow, precisely because the formal coherence of the response remains intact.

On this execution, NeoMundi therefore maintains a maximum coherence score (1.00) while simultaneously surfacing a factual hallucination signal (0.80) and a FLAG decision.

This case illustrates the value of multi-signal analysis: fluency and coherence alone could have been reassuring, while the factual signal makes it possible to isolate the fragile point within an otherwise credible response.

Weekly cartography

Observation Level: 🟢

This week, the cartography of AI Barometer #9 remains at the green level, but shows several localized movements.

Three profiles that were previously at 0% of flagged responses move above zero: PROFILE-212079, PROFILE-48C581 and PROFILE-739F7C each reach 0.25%.

In parallel, semantic variation increases on several profiles, notably PROFILE-DEA9C5 at 5.75% and PROFILE-F5FF91 at 4.01%. The dominant regime nevertheless remains NORMAL_SIGNAL for all 12 observed profiles.

The cartography therefore does not show a general shift of the panel, but several simultaneous displacements that become visible in the longitudinal follow-up.

Metadata of this execution: NM-WEEKLY-Q03 · repetition 97 · FLAG · coherence 1.0 · factual hallucination 0.8 · stability 0.676923.

NeoMundi weekly barometer #9

12
de-identified profiles
4,800
executions
99.729 %
coverage
92.28 %
average stability
2.09 %
semantic variation

NeoMundi Barometer · week 9 · 12 de-identified profiles · generated 15/08/2026 13:45

Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.

0 0.5 1 1.5 2 2.5 0 2 4 6 8 10 12 Responses needing more verification (%) the higher this number, the more responses need to be checked Change in response meaning (%) the meaning of responses changes from one run to another Meaning often changes, few responses need verification Meaning remains stable, more responses need verification Empty corner : no profile combines major meaning changes and more verification needs

Methodology and public data

This Barometer follows 12 de-identified profiles, each submitted to 4 fixed questions repeated 100 times. In total: 4,800 executions, of which 4,787 were fully scored, representing 99.729% coverage.

The published scores are measurement signals, not verdicts or rankings.

View the aggregated public data for Barometer #9

Leave a Comment

Your email address will not be published. Required fields are marked *


Scroll to Top