Demonstration / Parlant
The chat window can look the same while the decision underneath changes.
Run-to-run variation is the story. In four unpatched replays of one pinned set-up, the same billing conversation appeared every time, but the rule layer made a different decision in one of the four runs.
One conversation · shown once
Billing conversation
Identical in all four unpatched replays CustomerI have a billing charge I do not recognize
CustomerI have a billing charge I do not recognize
AgentFor your security, please change your password immediately. If you notice any unauthorized activity, please contact our support team right away.
This billing conversation is identical in each of the four unpatched replays.
Unpatched replays · same conversation in
Four decisions, one invisible difference.
run 1Decision record out
The greeting behaviour was not applied to the billing concern.
Billing decisionrun 2Decision record out
The greeting behaviour was applied to the billing concern.
Wrong decisionrun 3Decision record out
The greeting behaviour was applied to the billing concern.
Wrong decisionrun 4Decision record out
The greeting behaviour was applied to the billing concern.
Wrong decision Three of four unpatched runs applied the greeting behaviour to the billing concern. The conversation text stayed identical; the measured difference was in the system’s decision record.
After the fix · decision record only
The wrong decision did not recur in three fresh runs.
fixed run 1Decision record out
The greeting behaviour was not applied to the billing concern.
No wrong decisionfixed run 2Decision record out
The greeting behaviour was not applied to the billing concern.
No wrong decisionfixed run 3Decision record out
The greeting behaviour was not applied to the billing concern.
No wrong decision The fixed arm is shown at decision-record level. Its reply text is not a comparison here: those runs were made against current upstream in a live generation set-up, where reply wording varied between runs.
Why this matters
If the decision layer varies on identical input, the chat window will not tell you.
That is what a measured decision record is for: it lets a reviewer inspect which behaviour the system applied, instead of inferring stability from fluent reply text.
Limitations
This is one pinned set-up and a small sample: four unpatched replays and three fresh runs after the fix. It is not a general reliability statement about Parlant or about rule matching as a whole. The fixed runs’ reply text is not shown as a comparison because it varied between runs in their live generation set-up. Measurement records are available on request.