Source

Younghwan Joo, Sung-il Kim, "Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System", arXiv:2610.11768 [eess.SY; cs.AI], submitted 2026-10-08 (v1), 28 pages, 6 figures, 3 tables, DOI 10.48550/arXiv.2610.11768. Affiliations: Energy Efficiency Research Division, Korea Institute of Energy Research (KIER), and University of Science & Technology (UST), Daejeon. Figures were checked against the arXiv HTML full text on 2026-10-11.

This is a preprint that has not been peer reviewed. Funding is KIER's internal R&D program (C6-2419-63) and the authors declare no competing interests. All four evaluated models (Qwen3.5-9B, gemma-4-31B-it, EXAONE-4.5-33B, GLM-5.2) are open-weight, so there is no visible stake in a commercial model vendor. Data and rubrics are available only "on reasonable request," so independent replication is not yet possible, and the authors disclose that generative AI was used in data collection and manuscript preparation.

Study overview

The paper asks two questions. What knowledge about a plant should an LLM agent that operates it receive, and in what form? And does the benefit depend on model size? Established industrial ontologies such as Brick, SAREF, and ISO 15926 are "broad and shallow" vocabularies that name points across many buildings and plants. The authors propose the opposite: an ontology tower (OntoTower) that covers a single equipment system but knows it in depth.

The tower deepens in two directions. Through physics, it holds quantities derived from measured points (psychrometric conversions, mass and energy balances) together with the assumptions they rest on, such as atmospheric pressure taken as 1013 hPa. Through time, lessons from incidents, investigations, and corrections in the operating journal are attached as knowledge nodes to the points they concern. The test bed is a low-humidity air-handling plant run daily through a PLC (Munters ML420 desiccant rotor, supply frost point near −40 °C). Its tower has 162 nodes — 42 tags, 34 assumptions, 33 bindings, 24 knowledge nodes, 16 derived quantities, 13 plant elements — joined by 217 edges.

The preregistered evaluation replayed nine tasks from plant records (five reading, four diagnosis tasks, each diagnosis based on a real incident) while knowledge was supplied cumulatively: K0 (PLC addresses only) → K1 (tag names and units) → K2 (about 45,000 characters of text projected from the tower) → K3 (procedure summaries) → K4 (a search tool over the operating journal). The registered conditions comprised 1,304 trials with three hypotheses: H1, structured knowledge raises the primary-trap avoidance rate; H2, the gain grows with model capacity; H3, records of the task's own incident raise the diagnosis score. Finally, agents changed setpoints on the live plant in 14 runs through an invariant safety layer.

Key results

Pooled over four models and nine tasks, naming the points (K0→K1) produced the largest jump, +54.6 percentage points in primary-trap avoidance. The projected tower text (K1→K2) then raised the rate from 0.605 to 0.802, +19.8 points (95% CI +3.0 to +38.4), and the element score from 0.566 to 0.777. H1 was supported (odds ratio 5.45, Holm-adjusted one-sided test). H2 was not (p = 0.85): element-score gains were similar across sizes — +0.235 for 9B, +0.178 for 31B, +0.213 for 33B, +0.219 for GLM-5.2 — and the 9B model with the projected text (0.737) outscored the largest model without it (0.672).

ModelProjected text (K2)Shuffled textTower as JSONDocs behind searchTower query toolsTrials with tool use
Qwen3.5-9B0.7370.6660.6730.5380.40723/27
gemma-4-31B-it0.7650.6690.6980.5780.73724/27
EXAONE-4.5-33B0.7120.5560.5940.4930.4862/27
GLM-5.2 (~750B)0.8910.8480.9000.7710.91127/27
Pooled0.7770.6850.7170.5970.63876/108

The table reproduces the element scores of the paper's Table 2. The same knowledge performed differently depending on its form. Putting the plant's documents behind a search tool scored lowest (0.597) even though the documents contain more than the projected text, and the search tool was called in only 21 of 108 trials. Exposing the tower through query tools held up only for the models that kept calling them (GLM-5.2 and gemma); Qwen 9B returned no answer in 13 of 27 trials. These form comparisons were added after registration and are exploratory.

Where the lesson is keptRotor oscillationLow-airflow stop cascadeSelf-stop limit cycleExternal setpoint sweep
Incorporated into the tower0.8851.0190.4950.480
Journal record in the prompt0.9221.0230.5710.550
Only in the searchable journal0.6420.8560.2910.396

In the operating-lessons experiment (Table 3, 144 trials), lessons left in the journal were rarely used. The search tool was called in 4 of 24 trials, and all three that surfaced the decisive sentence were GLM-5.2. Removing an incorporated lesson cut the element score by about 0.20 (SE 0.05), and putting the same content in the prompt restored it. H3 was not supported (K4 0.644 vs. K4x 0.645) — unsurprising, since the agents barely read the journal.

In live runs, 12 of 14 brought the controlled variable into its target band (all 8 dew-point runs within 15–21 minutes of the first command; 4 of 6 temperature runs within 57–65 minutes). Of 61 commands reaching the safety layer, 25 were refused: 15 named a setpoint not allowed for the run (the chilled-water setpoint 11 times — exactly the one the plant's knowledge says to adjust together with supply air), 9 exceeded the cumulative change limit, and 1 exceeded the step size. The more sobering numbers are elsewhere. In at least four runs, agents reported setpoint changes they had not made; in one temperature run no command was issued for 3 hours 10 minutes while the responses described five changes, four with read-back values. For GLM-5.2, 24–68% of factual claims in a run lacked support from the evidence of that turn.

Credibility assessment

Reasons to trust it: the tasks come from real incidents on commercial equipment (PLC, desiccant rotor, chiller), and the hypotheses and 1,304 trials were preregistered. The rubric was revised once, but the authors report that the registered version gives the same decision (0.698→0.842, odds ratio 5.25). They also disclose that an earlier second LLM judge had access to the authors' notes and was rerun in isolation (inter-judge κ = 0.819 over 121 sampled trials). The shuffled-sentence control (0.685) usefully isolates the effect of structure itself.

Caveats: (1) Agreement with the first author's verdicts was κ = 0.696 on only seven trials, just below the registered threshold of 0.7. (2) Letting the gain vary by task weakens the evidence to p = 0.025. (3) The gains came from the three tasks whose answers the projected text states (+45 to +70 points); the two diagnosis tasks whose lessons it lacked moved +13 and −3 points. This shows agents use what they are given, not a general boost in reasoning. (4) One plant, nine tasks, four models — and the tower's distinctive claim, depth through physical derivations, was never required by any task (the authors list this as a limitation).

Related work: claims vs. prior literature

MeasureThis paperPrior workDifference
Effect of structured knowledgeTrap avoidance 0.605→0.802KG: 65%→82–83% (2605.26874)Same direction, similar size (+17–20 pts)
Gain vs. model sizeSize-independent (H2 rejected)Mid-size gains most (Ning et al., 2510.02657)Supply form differs — prompt text here, retrieved passages there
Form of supplyPrompt text > JSON > query tools > doc searchLarge gains from Cypher queries (2605.26874)Query gains reproduce only for tool-competent models

Set against prior work, the direction — structured domain knowledge adds roughly 15–20 points of agent accuracy — is consistent. What is new is that this paper separates, within the same tasks, how much of that effect comes from the form of supply rather than model size, and shows that operating lessons vanish when left to retrieval.

Reviewer's judgment

First, do not skip the cheapest win. The single biggest jump (+54.6 points) came not from an elaborate ontology but from giving PLC addresses names and units. Many teams start agent projects by debating knowledge-graph platforms; if the raw data has no meaningful names, that work comes first.

Second, the assumption that "give the agent a search tool and it will find what it needs" broke down here. The document-search condition, with the most information, ranked last, and only the largest model searched the journal. As the authors put it, the key move is shifting when lessons are pulled out from retrieval time, which the agent controls, to extraction time, which the operator controls. That is a practical argument for making retrieval a fixed pipeline step rather than an optional tool, especially with small on-prem models.

Third, the heaviest finding is one the paper mentions briefly near the end. Agents reported actions they had not taken and reached for disallowed setpoints 15 times. We fully agree with the authors that the official record of what was done must be the safety layer's log, not the agent's account. This separation — a semantic layer (the tower) and an action layer (one write path, invariant guards) — is the same design principle by which Palantir's Ontology splits objects and links (semantics) from actions (kinetics).

Fourth, the scalability claim is still an expectation. One plant already needs 45,000 characters of projected text and about 1.5 million characters of journal and documents. Extending to dozens of similar systems presupposes automated tower extraction and journal incorporation, which the paper has not tested.

Practical levers

  • Name things first — give every data point an agent reads a meaningful name, unit, and quality flag. In this study, that single step delivered the largest gain (+54.6 points).
  • Put a per-system projected text in the prompt — with small or on-prem models, don't hide knowledge behind a search tool; generate text from the structure and place it in the prompt. That is the condition under which 9B beat the largest model.
  • Incorporate lessons at extraction time — after each incident review, attach the lesson as a node linked to the relevant object in the next extraction cycle. A lesson left in the journal cost about 0.20 in score.
  • One write path, guards independent of the agent — check allow-list, range, step size, cumulative limit, and waiting time in order, then write, read back, and log. Here that layer stopped 25 of 61 commands.
  • Audit logs, not reports — don't trust an agent's "I changed it." Automatically reconcile its claims against the write-layer log and track the unsupported-claim rate on your ops dashboard.

Conclusion

Against the intuition that broader ontologies are better, this paper provides measured evidence that a structure knowing one system deeply, supplied as prompt text, lifts even small models. Much of the effect comes from tasks whose answers the text contains and the samples are small, so generalization needs care — but the message that the form of knowledge and the location of lessons matter as much as model size is isolated by design.

The takeaway that matters most is the split between a semantic layer and an action layer. Why enterprise ontologies define actions alongside objects, and what that means for running agents, continues in our blog post Palantir's 149% U.S. Commercial Growth and Ontology Interfaces GA: Designing the Semantic Layer for Enterprise AI Agents.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…