Demonstration / Parlant

The chat window can look the same while the decision underneath changes.

Run-to-run variation is the story. In four unpatched replays of one pinned set-up, the same billing conversation appeared every time, but the rule layer made a different decision in one of the four runs.

One conversation · shown once

Billing conversation

Identical in all four unpatched replays
Customer

I have a billing charge I do not recognize

Agent

Hi

Customer

I have a billing charge I do not recognize

Agent

For your security, please change your password immediately. If you notice any unauthorized activity, please contact our support team right away.

This billing conversation is identical in each of the four unpatched replays.

Unpatched replays · same conversation in

Four decisions, one invisible difference.

run 1

Decision record out

The greeting behaviour was not applied to the billing concern.

Billing decision
run 2

Decision record out

The greeting behaviour was applied to the billing concern.

Wrong decision
run 3

Decision record out

The greeting behaviour was applied to the billing concern.

Wrong decision
run 4

Decision record out

The greeting behaviour was applied to the billing concern.

Wrong decision

Three of four unpatched runs applied the greeting behaviour to the billing concern. The conversation text stayed identical; the measured difference was in the system’s decision record.

After the fix · decision record only

The wrong decision did not recur in three fresh runs.

fixed run 1

Decision record out

The greeting behaviour was not applied to the billing concern.

No wrong decision
fixed run 2

Decision record out

The greeting behaviour was not applied to the billing concern.

No wrong decision
fixed run 3

Decision record out

The greeting behaviour was not applied to the billing concern.

No wrong decision

The fixed arm is shown at decision-record level. Its reply text is not a comparison here: those runs were made against current upstream in a live generation set-up, where reply wording varied between runs.

Why this matters

If the decision layer varies on identical input, the chat window will not tell you.

That is what a measured decision record is for: it lets a reviewer inspect which behaviour the system applied, instead of inferring stability from fluent reply text.

Limitations

This is one pinned set-up and a small sample: four unpatched replays and three fresh runs after the fix. It is not a general reliability statement about Parlant or about rule matching as a whole. The fixed runs’ reply text is not shown as a comparison because it varied between runs in their live generation set-up. Measurement records are available on request.