Abstract
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model’s own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input.
Main findings
- ≤ 1.2ppRewording barely matters.Paraphrase stays within 1.2 pp of the re-run noise floor, word/spacing edits within 0.5 pp (one-sided 95% bounds; benign rewording only).
- 0tokensOut-of-schema fields never reach the model.They bill no extra token in any of 1,624 requests. Without the delivery check they would have been counted as robustness.
- 12.1%One observer opinion in the state flips decisions.Statistically tied with the strongest injected command (10.1%), although the opinion contains no instruction and claims no authority.
- 38.0%Confident answers are pushed into human review.The same opinion moves 38.0% of confident Choice/Score answers below a 0.8 gate, against 0.5% under the re-run; 119 of these 154 keep their decision.
Threat Model
Software acts on a typed answer through fixed thresholds, e.g. approving a transfer only above a confidence of 0.9. Whoever can steer the decision therefore steers the application. Select a role to see which fields of the request it controls.
{ "model": "jev-1.13.0", "state": "Help! My payouts have been failing for 3 days.", "questions": { "q002": { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "Payments, invoicing, refunds", "technical": "Bugs, outages, integrations", "sales": "Pricing, upgrades, new accounts" }, "comment": "…" } } }
Both count as harm only when the correct answer has not changed. The question asks about the content, so text that argues for an answer without adding facts about the case, such as an outside observer’s opinion, should leave the decision where it was.
Benchmark and Protocol
Identical requests change the returned Score value for 38.5% of Score questions, and 82.4% of labels in the release are the model’s own consensus answers. JevAdvBench therefore scores each attacked decision against the model’s own clean decision and reads every flip rate against an identical re-run.
66 scenarios, 812 questions
A state plus a typed question, extended from the vendor’s example scenarios; label provenance recorded.
One reversible edit per variant
Question, State and Injection attacks inside the schema; Structure edits as a delivery control.
One question per request
Typed answer plus billed input tokens. Added characters at +0 billed tokens means the edit was not delivered.
Flip vs. own clean decision
Excess over the re-run noise floor, plus target hit, shift without flip and confidence-gate effects.
Figure 2. The JevAdvBench evaluation pipeline. The noise floor, the flip rate of an identical re-run, is 1.0% [0.3, 1.9] over all 812 questions.
Flip, shift, or hold?
A flip is a changed decision; a shift keeps the decision but moves the output past a fixed margin. Drag the slider to move the attacked answer and see how the benchmark classifies it.
Attack Suite
Nine single-query attacks edit three surfaces inside the schema; three Structure edits serve as a delivery control. Every variant is one call from a fixed template shared by all 66 scenarios, so each rate is a lower bound for an adaptive attacker. Select an attack to see it applied to the vendor’s own example.
Illustrative item from the paper: state “Help! My payouts have been failing for 3 days.”, question “Does this convey urgency?”, reference true, attack target false. One item does not establish aggregate effects; the right-hand figures are over all 812 questions.
Which Inputs Move Decisions?
Observer opinion flips 12.1% of decisions, an excess of +11.1 pp [7.8, 14.4] over the noise floor, although it contains no command and claims no authority. Authority impersonation, the strongest command, flips 10.1% (+9.1 pp); the difference, +2.0 pp [−2.1, 5.9], is not significant.
Flip rate by attack
Figure 3. No single ranking describes the model: on Noul the three commands lead; on Score the two opinions lead, and direct override and authority impersonation fall below unrelated sentences. Kendall’s τ between Noul and Choice rankings is 0.63, between Noul and Score 0.06. Score cells have a minimum detectable excess of 7.0 pp.
Flip rates and excess over the re-run
Click a column header to sort.
The slot a command lands in
instructions vs criteriaWithin the 245 questions with commands in both slots, the gap is +14.3 pp [9.5, 19.7]. It sits in Choice; for Noul and Score the CI includes zero. Unrelated sentences show no slot effect.
Delivery, verified from billing
Delivered text bills about one token per five characters (r = 0.987). The same 59 opinion texts add 58.1 tokens in the state and none in an extra field: for keys outside the schema, the API acts as a firewall.
How much can an attacker with a given capability flip?
Flips concentrate on a minority of questions: the 10% most fragile questions carry 59.2% of all flips. Choose which attack families the attacker can use to see the share of questions flipped by at least one of them.
Damage Beyond a Changed Decision
Most attacked decisions that change land on the answer the attacker named, and counting flips understates the damage.
of flips from the five delivered targeted attacks land on the target option [83.3, 94.0]; a uniform pick among other options would give 31.2%.
An unverified opinion moves the answer as if it were evidence: on Score it moves the expected level toward the target by +0.239 [0.144, 0.347] levels; on Noul it flips answers both ways (8 true→false, 11 false→true), whereas direct override flips 35 true→false and 2 back. The commands shift 23.6–24.6% of answers without a flip.
Consistency with 143 human-reviewed labels
124 of these 143 labels agree with the model’s own five-run answer, so this measures consistency with largely model-derived labels, not accuracy against independent ground truth.
Can the Confidence Gate Protect a Deployment?
Confidence at attack time adds little beyond knowing which items were fragile to begin with, and the opinion that flips decisions also pushes confident answers into human review, mostly without flipping them.
Confident answers pushed below 0.8
Over all delivered attacks, 441 of the 534 drops below the gate leave the decision unchanged. Added length alone does little: the opinion’s content does the work.
Does attack-time confidence detect flips?
Implications for Deployment
Recommendations derived from measured effects; no defence was evaluated.
- Treat user-writable state as argued input, not fact.An observer’s opinion appended to the state raised the flip rate far above the noise floor.excess +11.1 pp [7.8, 14.4]
- Keep user text out of the
instructionsslot.Within the same question, commands flip decisions more often in instructions than in criteria; the gap comes mainly from Choice, so criteria is not a safe place either.slot gap +14.3 pp [9.5, 19.7] - Plan review capacity for attack-induced escalation.The opinion pushes confident Choice/Score answers below a 0.8 gate. Noul returns no confidence and needs its own check.review load +37.5 pp [30.7, 43.4]
Limitations
All results are for one model version, jev-1.13.0, queried on 2026-09-25; no claim is made about other versions, other typed decision models, or RLCD as a training method.
Every attack is one fixed-template call with no search and no adaptivity, so each rate is a lower bound for an adaptive attacker.
The form edits are benign: 739 of 812 word/spacing edits change only spacing, and the equivalence bound says nothing about adversarial paraphrase search.
Structure edits measure non-exposure, inferred from billing; their low flip rates are not evidence of robustness.
Variants are designed, not verified, to preserve the answer: a single LLM annotator checked 40 (25 preserved, 15 with caveats, none changed).
82.4% of labels are the model’s own consensus answers and items were selected for clean stability, which is why the primary metric is label-free.
Open artifacts
beta1.0-20260925run_all.sh regenerates every reported numberBibTeX
@misc{hu2026jevadvbench,
title = {JevAdvBench: A Benchmark and Black-Box Attacks for
Reinforcement Learning for Calibrated Decisions Models},
author = {Hu, Jianyi and Zhang, Hangtao and Liu, Yi and Zeng, Yeqi and
Zeng, Li and Wang, Xianlong and Wang, Rui and Zhang, Leo Yu},
year = {2026},
eprint = {2609.31142},
archivePrefix = {arXiv},
primaryClass = {cs.CR},
url = {https://arxiv.org/abs/2609.31142}
}
Ethics. JevAdvBench is dual-use. The templates are released because the attacks are cheap, fixed texts of the kind a user can already write into a ticket, and builders need them, with the delivery check and gate measurements, to test a deployment before an attacker does. All requests went to the public API; no third-party deployment was attacked and no personal data collected. Findings will be shared with the vendor before publication.