Measure Parlant’s own behaviour
Identical conversations were rerun to expose run-to-run variation and a known false guideline activation. The measured fix is now a public pull request.
Integrations / Parlant
These are two different pieces of work. The first measures Parlant’s own guideline selection. The second replays a normalized Parlant activation through Forge’s existing coaching arbitration. Neither is presented as a broad reliability guarantee or a live production connector.
Identical conversations were rerun to expose run-to-run variation and a known false guideline activation. The measured fix is now a public pull request.
One integration test loads normalized activations from a frozen Parlant run. Two separate developer-authored controls exercise the same adapter and existing deterministic coaching pipeline. The adapter has no matching or safety decision logic of its own.
What this does and does not show
The evidence is limited to the documented test set-up, one frozen fixture path and two synthetic control cases. It shows that those records can pass through Forge arbitration without changing that arbitration; it does not establish a live Parlant connector.