NeoMundi Observatory ran a full runtime protocol on Kimi K3 (Moonshot):
1,200 valid executions – 120 situations repeated 10 times across 12 families (factuality, reasoning, governance, medical, legal, cybersecurity, business operations, and more).
This is not a ranking.
This is not a judgment.
This is a runtime observation.
The objective is to make visible what happens at the exact moment a model generates a response: stability, coherence, semantic variation, factual signals, latency, volume, density, and the complete evidence chain behind every signal.
It is to show what a runtime metrology layer can observe at the very moment an AI system generates a response: its stability, coherence, variations, factual signals, latency, generation volume, and associated trace.
| Protocol | Result |
|---|---|
| Executions analyzed | 1,200 |
| Final coverage | 100% |
| Distinct situations | 120 |
| Repetitions per situation | 10 |
| Use-case families | 12 |
1. One model, multiple runtime regimes
A model does not behave the same way on a factual question, a reasoning task, an ambiguous prompt, or a complex business request.
Across the 12 families we observed clear differences in:
- latency
- response length
- information density
- stability patterns
A global average creates the illusion of uniformity.
Runtime metrology shows the opposite: the same model can operate under several distinct regimes.
For any organization the real question is no longer “Does the model work?”
It becomes:
Under which conditions does its behavior change, and are those changes acceptable for the intended use?
2. Stability is not truth
Stability measures the ability to produce comparable outputs when the same situation is repeated.
But a highly stable answer can still be incorrect, incomplete, or logically fragile.
🔔 Stability does not equal truth.
A model can reproduce the exact same error with remarkable consistency.
Conversely, variation across repetitions is not automatically a defect, it can be legitimate reformulation or a sign of inherent ambiguity.
Measurement should help distinguish between:
- normal variation;
- semantically significant variation;
- a stable but potentially incorrect response;
- a regime shift that deserves investigation
3. Performance is more than speed
Latency is useful, but it is only one dimension.
True performance combines:
- response time;
- token volume;
- information density;
- stability;
- coherence;
- factual or semantic signals;
- runtime cost and effort.
Measuring latency alone means watching the infrastructure while remaining blind to the generation itself.
Runtime metrology connects the two.
4. A signal only has value if it can be reconstructed
An aggregated metric can raise attention.
✔️ To become actionable, it must be traceable to a precise observation.
Every measurement in this study is linked to a complete evidence chain:
- prompt and family
- repetition index
- requested and observed model
- full generated response
- calculated metrics
- timestamp
- request_id and trace_id
- runtime decision + detailed reasons
When a signal can be reconstructed, it becomes operational.
It can support investigation, longitudinal comparison, human escalation, rerouting, or a governance rule.
What this observation demonstrates
- A model does not have a single behavior, it moves across different regimes.
- Stability must always be read with factuality, coherence, and semantic variation.
- Performance must combine speed, volume, density, cost, and behavioral quality.
- Every signal must remain linked to reconstructible runtime evidence.
🎯 NeoMundi does not add another dashboard.
It provides the missing measurement layer between model behavior and operational decision-making.
A measured signal is not a verdict.
It is a new capacity to observe, compare, investigate, and govern.
Download the full technical report
Runtime Observation Report
(Free PDF report)

NeoMundi Observatory · Independent Runtime Metrology
Analysis of 1,200 executions across 12 situation families
Moonshot Kimi K3 – 2026-07-22

Got to say even before I take the deep dive this cannot be disputed: “stability does NOT = Truth”
Excellent
Thanks!
Yes, « stability does NOT = Truth » is spot on and one of the key points.
Looking forward to your thoughts after the deep dive. 🤝