Observation Level: 🟡 – This week, AI Barometer #6 explores two complementary dimensions.
On the one hand, a form of volume variability: the same system can reach the same correct result while generating very different quantities of tokens.
On the other hand, a return to calm after the peak observed the previous week, with a global clustering of the profiles and the disappearance of the main factual tension signals.
Same correct result, very different token volumes
For the same question, the same system can reach the same correct result while mobilizing very different generation volumes. This still under-documented form of variability is one of the main focuses of this edition.
The question
We focused on Question 1 of the protocol:
“A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?”
Here are two responses produced by the same profile, both leading to the correct result.
Response A – 151 measured tokens
“Let x be the price of the ball in euros.
The bat then costs x + 1 €.
The equation is:
x + (x + 1) = 1.10
2x + 1 = 1.10
2x = 0.10
x = 0.05
The ball therefore costs 5 euro cents.”
Response B – 395 measured tokens
“To solve this problem, let us follow these steps.
Let us define the variables:
Let x be the price of the ball in euros.
The price of the bat is therefore x + 1 euro, since it costs 1 € more than the ball.
Let us write the equation:
The total cost of the bat and the ball is 1.10 €.
x + (x + 1) = 1.10
Let us solve the equation:
x + x + 1 = 1.10
2x + 1 = 1.10
2x = 1.10 – 1
2x = 0.10
x = 0.10 / 2
x = 0.05
Conclusion:
The ball costs 0.05 €, or 5 cents.
The bat then costs 0.05 € + 1 €, or 1.05 €.
Verification:
1.05 € for the bat + 0.05 € for the ball = 1.10 €.
The result is therefore correct.
Final answer: The ball costs 0.05 €.”
What this reveals
Both responses are correct and reach exactly the same result. Yet the second one mobilizes approximately 2.6 times more measured tokens than the first, representing an increase of about 162%.
The longer response adds a structured presentation, a verification step, and several reformulations. These elements may be useful depending on the context, but they do not alter the fundamental reasoning or the conclusion reached.
Potential consequences
This variation can have direct consequences on generation latency, large-scale usage costs, information density, the regularity of the user experience, and the predictability of the resources mobilized.
Conclusion
Result stability does not guarantee runtime footprint stability.
Measuring tokens in parallel with content makes it possible to distinguish a genuinely enriched explanation from a volume increase that provides only limited informational gain.
Return to calm after the peak observed the previous week
Observation Level: 🟡 Yellow
🧭 The weekly Barometer shows a clear improvement compared to the previous edition.
While several profiles showed a marked increase last week in responses requiring additional verification, this week’s cartography shows an almost complete clustering of the cloud into the main zone. The factual tension signal has strongly receded. Some semantic variations remain visible, however, which leads NeoMundi to classify this week as 🟡 Yellow: clear improvement, continued monitoring, with no recommended change in usage methods.
The visual comparison between the two campaigns shows a clear shift in the observed regime: Barometer #5 indicated an unusual collective tension, while the present edition points to a return toward more stable and better-controlled operation.
Mapping of 12 de-identified observed systems. Horizontal axis: responses needing more verification. Vertical axis: change in response meaning. Graphic specification comparable across all weeks.
The main movement this week is therefore a return to normal. The episode observed the previous week does not appear to be continuing at the same intensity. Out of methodological caution, the Observatory nevertheless maintains a 🟡 Yellow reading, pending confirmation in the next campaign.
Status of the surveillance protocol - profile DEA9C5
Profile DEA9C5 had experienced a marked break, moving from an initial regime close to 0.3% factual risk to 12.3%, then to a confirmed plateau around 14% across several follow-up runs.
After several weeks of persistence, the measurements from Barometer #6 indicate a clear return to calm, with factual risk brought down to 0.03%, stability rising to 92.30%, and semantic variation decreasing to 4.79%. This improvement nevertheless remains to be confirmed over time before concluding a full return to the initial regime.
| Indicator | Barometer #5 | Barometer #6 | Evolution |
|---|---|---|---|
| Average factual risk | 3.95% | 0.03% | strong decrease |
| Semantic variation | 8.05% | 4.79% | decrease of 3.27 points |
| Average stability | 80.78% | 92.30% | +11.52 points |
| Coverage | 96.25% | 99.25% | +3 points |
| Rate of flagged responses | 4.00% | 0.00% | disappearance this week |
| Average coherence | 77.66% | 100% | clear improvement |
Follow-up of a punctual signal - profile B912EA
Profile B912EA showed a punctual spike in Barometer #5, with average factual risk rising from 0.15% to 7.00% and 9.75% of flagged responses. In Barometer #6, both indicators return to 0.00%, while stability and coherence improve significantly. The signal therefore does not appear to have become established over time, but it justifies passive monitoring to verify that it does not reappear in upcoming campaigns.
| Indicator | Barometer #5 | Barometer #6 | Evolution |
|---|---|---|---|
| Average factual risk | 7.00 % | 0.00 % | disappearance of the signal this week |
| Rate of flagged responses | 9.75 % | 0.00 % | strong decrease |
| Average stability | 78.71 % | 92.31 % | +13.60 points |
| Average coherence | 75.20 % | 100 % | clear improvement |
| Semantic variation | 0.50 % | 0.50 % | stable |
| Coverage | 99.25 % | 99.25 % | stable |
Methodology and public data
This Barometer tracks 12 de-identified profiles, each tested against 4 fixed questions repeated 100 times. In total: 4,800 executions, with 99.438% coverage.
The scores are measurement signals, not verdicts or rankings.
