HEXIS Compiles LLM Agent Skills into Deterministic State Machines
A September 2026 paper introduces HEXIS, a system that takes a natural-language agent skill document and compiles it into an extended finite state machine. Instead of passing the full skill as context and asking the model to infer which step comes next, HEXIS makes control flow explicit and deterministic, with the language model invoked only inside states where semantic reasoning is genuinely required. The reported result: average task success improves by 16.1 percentage points over Skill + ReAct across four benchmarks and four executor models, with Qwen3.8-27B reducing token use by 38.4 to 88.9%.
The Problem with Native Skill Execution
Agent skill documents package domain knowledge, step-by-step instructions, tool interfaces, and worked examples into a reusable unit. The approach is broadly useful: according to SkillsBench (Li et al., 2026), curated skills increase the average pass rate from 33.9% to 50.5% across 87 tasks in eight domains. The problem is how agents execute those skills at runtime.
In standard practice, the skill document is supplied as context alongside the task input and interaction history. At each step, the model selects the next action through a distribution over the full skill, history, and current state: at = π(· | D, x, ht). The model must simultaneously apply task knowledge and recover its position within the skill’s procedural requirements, check whether a condition is met, identify which operation comes next, and then execute it. This coupling of knowledge application with control inference is the source of compliance failures.
The authors, WorldBuilder013 and Minghao Li (arXiv:2609.30123, September 24, 2026), formalize this as a local deviation probability at each step. Even with deterministic decoding, a model may correctly identify that a check failed and still choose to report rather than revise, because reporting is more likely under the distribution conditioned on the skill document and history. “Correct evidence alone does not ensure εt(ht) = 0; applying the control requirement remains a model decision, even with deterministic decoding.” Compliance failure probability compounds across steps in a long task: full compliance probability across K steps equals the product of per-step compliance probabilities, so even a 5% per-step error rate produces roughly 60% failure over 10 steps.
AgentIF, SOPBench, and prior skill-compliance research confirm this picture. The model knows the skill requirements; it simply does not reliably follow them, especially in longer tasks where the growing interaction history dilutes the relevance of early instructions.
How HEXIS Works
HEXIS takes a skill document, tool specifications, a task input schema, and development traces as inputs. It produces an extended finite state machine (EFSM) that encodes control requirements explicitly while keeping task knowledge inside model-accessible state prompts.
The compiled EFSM has the form M = (Q, q0, V, a, E, F, τ), where Q is a finite set of states, q0 is the initial state, V is a set of typed variables, a specifies each state’s operation and configuration, E specifies ordered outgoing transition edges with guard conditions, F is the set of terminal states, and τ specifies terminal outcome categories. Within each state, a local model prompt contains the subset of the skill’s knowledge relevant to that operation. Typed variables store intermediate values across states.
When the machine transitions, it evaluates guard conditions on the current variable values rather than asking the model to decide. Following a successful nonterminal operation, the runtime selects the first enabled outgoing edge: the next state is determined by the first guard condition that evaluates true. A model may write a verdict based on skill criteria; the runtime routes a failed verdict to revision through the guard condition without another model decision. This removes the opportunity for control deviation at that boundary.
The compiler constructs the machine in two stages. The initial machine is built from the skill document by mapping skill clauses and tool interfaces to state operations, local instructions, data bindings, and transitions. Then, the compiler aligns development traces with the existing states. Where a trace reveals a missing operation or dependency, the compiler adds or reuses states and refines transitions. Crucially, updates are accepted only after passing static checks and successfully replaying both the new trace and all previously accepted traces. This makes the machine verifiable against observed behavior before deployment.
Results: What They Measured
The authors evaluated HEXIS across four benchmarks with four executor models: qwen3.6-flash, Qwen3.8-27B, and two others. The state machines were compiled using qwen3.6-flash executions and then transferred unchanged to the other three executors.
On average across all four benchmarks and four executors, HEXIS improves task success by 16.1 percentage points over the Skill + ReAct baseline. The compiled machines lead or tie the baseline in 11 of 16 settings and outperform it in 15 of 16. The one setting where HEXIS does not improve provides the useful signal that some tasks, those whose workflow genuinely changes continuously with task input, may not map cleanly to a fixed EFSM.
The token efficiency results are substantial. Qwen3.8-27B reduces execution tokens by 38.4% to 88.9% depending on the benchmark. The range is large because token savings scale with how much of the original execution was spent on the model inferring its position and deciding next steps. Tasks with highly sequential, predictable control structures see the largest reductions.
On SpreadsheetBench, combining HEXIS state machine execution with SkillOpt (a method that optimizes the skill document using execution feedback) reaches 84.2% success. This combination is interesting because it addresses both sides of the problem: SkillOpt improves the knowledge content fed into states, while HEXIS handles the control structure. The combination is additive rather than substitutive.
The transfer result is particularly notable. State machines compiled from one executor model’s traces work for other models without recompilation. This suggests the compiled machine captures a property of the skill and its execution requirements, not a property of the particular model’s behavior.
Limitations and Open Questions
The compiler must correctly infer state boundaries, guard conditions, and data dependencies from natural-language skill documents and execution traces. The paper validates this on four benchmarks; the question is how well the compiler generalizes to skill documents with ambiguous or underspecified procedural requirements. Natural-language skill documents can be inconsistent, incomplete, or context-dependent in ways that are hard to compile reliably.
The setting also assumes that the optimal workflow for a skill can be represented as an EFSM with a finite state set. For tasks whose correct action sequence depends continuously on new information that cannot be pre-specified as a guard condition, this representation may lose information. The one setting where HEXIS underperforms the baseline likely represents this failure mode.
The token reduction numbers are author-reported results from four benchmarks. Independent evaluation on longer-horizon tasks with more complex skills, and at larger model scales, would strengthen the quantitative picture. The claim that 84.2% on SpreadsheetBench represents a meaningful improvement depends on what the baseline achieves on that benchmark specifically, which the abstract reports as an improvement but does not give the baseline absolute number for direct comparison.
Finally, trace-guided compilation requires development traces, meaning the system needs some prior executions to guide the compiler. Cold-start compilation from a skill document alone produces an initial machine, but the improvement from trace alignment is part of what drives performance. For genuinely new skills with no prior execution history, the compiled machine may be less reliable.
What This Means for Engineering Teams
The architectural shift HEXIS represents is worth paying attention to for any team building production AI agents. The dominant pattern in current agent frameworks is to pass a skill or system prompt as context and let the model drive control decisions at each step. HEXIS demonstrates that separating knowledge from control flow produces measurable gains in both compliance and efficiency.
For teams using ReAct-based agents on multi-step tasks, the token reduction alone may justify compilation. A 38-to-89% reduction in execution tokens is significant in both cost and latency terms, particularly for Qwen3-class models running on production infrastructure. The compilation overhead is paid once; the savings apply to every execution.
The incremental compiler design, where updates are accepted only after replay passes, is also relevant beyond HEXIS specifically. It is a form of regression testing applied to control flow. Teams building any kind of structured agent workflow can adopt this principle: changes to a workflow specification are not accepted unless they reproduce the expected behavior on all previously validated traces. This connects to broader work on optimizing agent inference stacks and making agent behavior deterministic where possible.
The combination with SkillOpt points toward a practical development pipeline: optimize the skill document with execution feedback, then compile the optimized document into an EFSM for deployment. The skill optimization phase can use natural-language feedback; the compilation phase produces the artifact that actually runs. Teams working on ReAct agent architectures should watch whether this pattern becomes a recommended baseline for skill-based agents over the coming months, particularly as agent tasks grow in length and complexity.
Key Takeaways
- HEXIS compiles agent skill documents into extended finite state machines, with skill knowledge in state prompts and control requirements encoded as guard conditions on transitions.
- Average task success improves by 16.1 percentage points over Skill + ReAct across four benchmarks and four executor models; the machines were compiled from one model’s traces and transferred unchanged to others.
- Qwen3.8-27B reduces execution tokens by 38.4 to 88.9% depending on the benchmark, because the model no longer needs to repeatedly infer its position in the skill and select the next step.
- Combining HEXIS with SkillOpt (skill document optimization) reaches 84.2% on SpreadsheetBench, showing the two approaches address different problems and combine additively.
- The incremental compiler accepts updates only after passing static checks and replaying all previously accepted traces, providing a built-in regression guarantee.
- Tasks whose optimal workflow genuinely varies continuously with new input may not map well to a fixed EFSM; the one setting where HEXIS underperforms suggests this boundary exists.
Work With Origins AI
Origins AI builds production AI systems for engineering teams. If your team runs multi-step agents on business workflows and compliance failures or inference costs are a real problem, talk to our team.

