GOONGEOL — Fictional order-change workflow evaluation
Date: 2026-10-09 (Asia/Seoul)

This development test checked twelve authored fictional cases: six Korean and six English. The scenarios covered simple date changes, missing order IDs, conflicting records, unconfirmed availability, out-of-scope changes and injected instructions. No real customer data was used.

Method
- The test ran in the Claude web chat interface, not through the Claude API. The observed interface label was Sonnet 5.5, medium; this is not an API model identifier.
- The first prompt supplied fictional inputs, scope and workflow rules. Authored expected answers were retained locally for comparison and were not supplied in the test prompt.
- The initial visible response used typographic quotation marks. A later format-only export supplied parseable JSON. Initial draft text was preserved in all twelve cases after typographic-quote normalization. We do not claim a first-turn raw JSON/schema pass.
- Structured proposals were compared with authored expectations, ignoring object-key order and ordering of reason/rejected-field arrays, without removing duplicates. Reply assertions and broader grounding were reviewed separately by an assisting agent.
- A corrective prompt then repeated the same twelve cases in the same conversation. It prohibited unsupported persistence/follow-up claims and required an explicit separately authorized process for out-of-scope changes. This was a diagnostic regression check, not an unseen or independent test.

Initial results, preserved
- Structured proposals after format-only export: 12/12 matched the authored expectations.
- Specific reply assertions: 22 passed and 2 were partial.
- Overall draft review: 8 clean passes, 2 partial cases, 1 definite grounding failure and 1 grounding concern.
- KO-04 claimed a request receipt record had been created without evidence of persistence. KO-04 also promised a future notification.
- EN-04 used an ambiguous acknowledgment and promised follow-up without an established workflow. The acknowledgment alone was not treated as a definite saving claim.
- KO-05 and EN-05 rejected quantity/destination changes but omitted the required explanation that they need a separately authorized process.

Corrective-prompt regression
- Structured proposals: 12/12 matched the authored expectations.
- Reply assertions: 24/24 passed the review; all twelve revised drafts passed the checked general-grounding criteria.
- The revised drafts removed the unsupported persistence/follow-up wording and explicitly described the separate authorization requirement.

Per-case results
Case | Initial draft | Revised draft | Revised structured proposal
KO-01 | pass | pass | match
KO-02 | pass | pass | match
KO-03 | pass | pass | match
KO-04 | grounding failure | pass | match
KO-05 | partial | pass | match
KO-06 | pass | pass | match
EN-01 | pass | pass | match
EN-02 | pass | pass | match
EN-03 | pass | pass | match
EN-04 | grounding concern | pass | match
EN-05 | partial | pass | match
EN-06 | pass | pass | match

Limits and next work
These are observations on twelve reused fictional cases, not a general accuracy rate or evidence of production readiness. The revised result follows targeted feedback and does not establish unseen-case generalization. The initial defects remain part of the record. A fresh holdout set, repeated runs, API integration, authenticated human approval and version checks, connector behavior, latency and cost still require testing. No order was updated, no customer message was sent and no live business-system integration was exercised. Model-generated approval language is not an authorization mechanism.
