# Evolution 3 — limitations (preserved)

1. **Rule models, not learned models.** The exec model is a hand-written table. The benchmark found
   precision 0.71 and recall 0.92 for "modifies the filesystem" on 35 real commands; `find -fprint` is the known
   miss. Treat it as a consequence model, not a command-failure model.
2. **Adverse-outcome calibration is worse than a base-rate predictor** (Brier 0.37 vs 0.16 on the benchmark).
   Because history only raises a forecast, over-prediction is never corrected automatically. Lowering a forecast
   needs a governed model change (a new version, which invalidates in-flight forecasts). This is a deliberate safety
   trade-off ("historical success cannot create authority"), and it costs calibration.
3. **Observation is shallow.** Exec effects are observed only in the contract's cwd (depth 2, at most 2,000
   entries; otherwise UNOBSERVABLE). Network, process and out-of-cwd effects are not observed.
4. **Memory readers are authorized readers, not actual readers.** MemVault classifier recall is 0.617 on its
   held-out set.
5. **No review workflow.** Every "require more evidence" outcome maps to DENY (fail closed). There is no human
   approval state yet (roadmap #12).
6. **Many dimensions are NOT_MODELED** (financial, economic, reputational, social, organizational, supply_chain,
   identity, model) for the four products. Physical is NOT_APPLICABLE: there is no actuation substrate.
7. **Drift needs volume:** 100 observations per (kind, effect class) before it can fire.
8. **Single host.** The corpus, budget and policy live in the tenant SQLite file shared by blue and green. There is
   no cross-host replication, and the quorum stays in-process (see E2 LIMITATIONS).
9. **The prompt's numeric targets were not met:** 150 invariants, 1,000+ adversarial tests and 500+ per category.
   What exists: 18 named invariants backed by tests, 45 E3 tests, 3,000 property cases, and 1,500 generated attack
   iterations across 6 classes. No padding was added.
10. **Scope:** 4 of the 20 Frontier products (owner directive). The other 16 have no predictive integration.
