A benchmark for industrial control · Preprint 2026

PLCWorldBenchmarking LLM-Generated PLC Programs
in Closed-Loop Plant Simulation

Yunji Kim1,*Yunseok Lee1,*Hyunwoo Seo2Jaerim Choi2Woojin Lee1,†

1 Dongguk University2 MOAI technologies

* Equal contribution   ·   † Corresponding author

Inside PLCWorld Shared robot & machining cells

From code to consequences

What happens
after the program runs?

Follow each control command through device motion, workpiece transfer, and sensor feedback.

Explore the execution loop
PLC programPlant
100
Synthetic tasks
473
Task–condition pairs
2
Task suites
3
Dependency scopes

01 / Overview

Evaluate programs through plant execution.

Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. To assess an LLM-generated PLC program, we need to observe how those commands affect device and workpiece states.

PLCWorld couples Structured Text (ST) execution with simulated plant responses and sensor feedback in a common closed-loop environment. It provides 100 synthetic tasks and 473 registered task–condition pairs across Motion Control and Material Handling, grounded in control relations identified in industrial PLC programs and engineering documentation.

Our protocol measures Task Success and Safety Violation separately. Comparing these outcomes with submission-profile acceptance reveals the Execution Gap: programs that pass an upstream check can still fail their task or violate a constraint during execution.

PLCWorld framework: task specifications guide ST generation; a PLC runtime and simulated plant exchange control outputs and sensor feedback; execution records determine task success, safety violations, and the execution gap.
The PLCWorld framework. Generated programs interact with the plant through repeated PLC scans.

Task Success

Reach the required goal state with the specified events, ordering, and hold conditions.

Safety Violation

Track breaches of specified constraints throughout execution, including after completion claims.

Execution Gap

Measure task failure or observed violation among cases whose programs pass the upstream check.

02 / Tasks in motion

Two suites. Three scopes of control.

Six representative tasks from Figure 3 of the paper, recorded in the PLCWorld replay viewer.

Easy

Within one function

Coordinate a basic control operation.

15 tasks per suite
Medium

Between functions

Connect the dependent functions of one job.

20 tasks per suite
Hard

Between jobs

Coordinate independent jobs sharing resources.

15 tasks per suite

03 / Evaluation

Completion and safety tell different stories.

Six LLMs and four adapted generation-and-verification workflows, evaluated under a common protocol.

Direct GPT-5.5 · Pooled by case count

Performance drops as the scope of coordination grows.

Easy82.70%
Medium84.09%
Hard25.10%

Task Success across both suites. These comparisons describe different task groups; they do not isolate the effect of dependency scope.

Paper Figure 5 compares suite-combined success rates by difficulty for direct models and adapted workflows.
Success rate by difficulty, as reported in the paper.
Direct ST generation · Motion Control · Results (%)
Model / MethodEasyMediumHardExecution
Gap ↓
SR ↑SVR ↓SR ↑SVR ↓SR ↑SVR ↓

SR: Task Success rate. SVR: Safety Violation rate among cases entering plant execution. Execution Gap is pooled across suites and conditional on submission-profile acceptance. Direct rows pool three independent generations per task; workflow comparisons use a matched GPT-5.5 baseline.

Validation beyond the reference program

Practitioner review, alternative programs, specification–evaluator alignment checks, and independent ST-runtime comparisons. All 542 targeted counterexamples activate their designated rules under at least one registered condition.

04 / The execution gap

A command is only the beginning.

The same accepted program can produce different outcomes as device response times change.

An ST program closes a gripper and waits 500 milliseconds before lifting. A faster grasp succeeds; a slower grasp leaves the workpiece unheld, so lifting fails despite passing the upstream acceptance check.
Figure 1. A fixed delay can elapse before a workpiece is securely held. Plant execution exposes the consequence.

05 / Citation

Build on PLCWorld.

If you use PLCWorld in your research, please cite our preprint.

BibTeX
@misc{kim2026plcworld,
  title  = {PLCWorld: Benchmarking LLM-Generated PLC Programs
            in Closed-Loop Plant Simulation},
  author = {Kim, Yunji and Lee, Yunseok and Seo, Hyunwoo
            and Choi, Jaerim and Lee, Woojin},
  year   = {2026},
  note   = {Preprint}
}