What a Stable Average Can Conceal in Production AI

James Moore’s analysis of AI Barometers #6 and #7 shows how a minimal global shift can mask concentrated changes across specific profiles, tasks and alert types.

NeoMundi’s AI Barometers were designed to make the behaviour of AI systems observable over time. Once the measurement is produced, however, another question becomes essential: what can an apparently stable average conceal?

In his fourth contribution to NeoMundi Research, James Moore, a contributor specialised in human governance at execution time, examines the public aggregate results of AI Barometers #6 and #7. His work starts from a simple observation: looking only at the overall mean or the dominant regime of a campaign can cause important parts of what is actually happening inside the observed system to be missed.

Between AI Barometers #6 and #7, the mean stability changes very little: it moves from 0.922864 to 0.920596, a difference of just 0.002268. Taken in isolation, this movement could suggest an almost unchanged panel. When Moore looks beneath the average, however, the structure appears differently.

Fewer exceptions, but different exceptions

The total number of complete non-normal observations decreases between the two campaigns, falling from 131 to 106. One might quickly conclude that the situation is improving. Yet their composition changes markedly.

  • Observations showing only semantic variation drop from 113 to 62.
  • Factual-only alerts rise from 16 to 31.
  • Combined factual and semantic alerts rise from 2 to 13.

Overall, observations involving a factual dimension therefore increase from 18 to 44.

The number of exceptions decreases, but those that remain change in nature. This is one of the central findings of the report: the volume of a signal is not enough to describe its meaning. Its composition also matters.

An average can conceal a localised event

The small movement in global stability also masks a significant concentration.

PROFILE-212079 moves from a mean stability of 0.923077 to 0.895175. Its FLAG rate simultaneously rises from zero to 9 %, while the other profiles remain close to their previous stability range.

In other words, the slight decline observed at campaign level is not necessarily a sign of general degradation. It is largely concentrated in one particular profile.

The distinction is fundamental. A diffuse variation across an entire panel and a localised rupture that produce the same average do not tell the same story. The distribution matters as much as the mean.

Certain tasks also concentrate certain signals

Moore also observes concentration at the question level.

Question 4 accounts for approximately 95 % of all semantic-positive observations in Barometer #6 and still around 80 % in Barometer #7. At the same time, in Barometer #7, Question 2 becomes the question with the highest FLAG rate and the lowest mean stability.

At the same time, in Barometer #7, Question 2 becomes the question with the highest FLAG rate and the lowest mean stability.

This suggests that the same system can produce different types of signals depending on the nature of the task it faces.

It is therefore not enough to know that an anomaly exists. Several further questions must also be answerable:

  • Where is it concentrated?
  • Does it recur over time?
  • On which type of task?
  • Which signals led to its classification?

From detection to structured evidence

This is where Moore’s contribution goes beyond descriptive analysis.

He identifies six dimensions that need to be strengthened: concentration, recurrence, materiality, task context, alert lineage and the hand-off to governance. He proposes, in particular, the distinction of several longitudinal states: isolated, recurrent, persistent, expanding or resolved.

These states would not be verdicts. They would simply describe the observable history of a signal over time. The distinction is essential.

An anomaly observed only once does not carry the same weight as a phenomenon that reappears week after week. A variation concentrated on a single task is not equivalent to a movement that gradually spreads across multiple profiles.

Measuring without deciding

This approach respects a central boundary of NeoMundi: measurement produces evidence; it does not carry decision-making authority.

James Moore formulates this separation as a chain:

Signal → interpretation → policy → accountable action.

NeoMundi can make changes observable and document their concentration, recurrence and lineage. It is the organisation operating the system that must determine their materiality, operational impact and any decision to be taken, within the framework of AI systems governance.

This is precisely where longitudinal measurement takes on its full meaning. An average can say that a system appears globally stable. The distribution shows where one must start looking.

The full contribution by James Moore is offered within the NeoMundi Research / Expert Lens framework and is submitted for methodological review.

Download the complete analysis by James Moore

James Moore

Leave a Comment

Your email address will not be published. Required fields are marked *


Scroll to Top