source: arxiv artificial intelligence: ontology-amplified distillation and contextuality auditing for sovereign enterprise language models: a combined proof-of-mechanism and negative-results method study

level: research

regulated financial institutions with data-residency rules need language models that run inside their own systems. this paper combines two studies into one article. the first part is a proof-of-mechanism study of ontology-amplified distillation. a qwen3.6-27b student model is adapted to the foundation agenticos ontology. it uses supervised fine-tuning on frontier-teacher trajectories and ontology-grounded direct preference optimization. training happens locally on a single apple m5 max with 47 synthetic english preference pairs from different domains.

the distilled student is tested on 40 held-out vietnamese financial tasks. it grounds 36 of 40 tasks, giving a grounded rate of 0.90. the mean ontology term-coverage is 0.95 on a metric floored at 0.50. this matches the gpt-5 frontier baseline, which also grounds 36 of 40 tasks. however, the study is underpowered to establish equivalence. the small sample size limits strong conclusions.

the second part is a negative-results study on contextuality auditing. it examines whether the model's reasoning stays consistent with the ontology under different contexts. the findings show no significant improvement from the auditing method. the combined work highlights challenges in verifying ontology alignment for small-scale, locally trained models. it suggests that more data and larger studies are needed to confirm the distillation approach.

why it matters: it shows a potential way to build compliant, in-house ai for regulated industries, but the small scale means the method needs more testing before real-world use.


source: arxiv artificial intelligence: ontology-amplified distillation and contextuality auditing for sovereign enterprise language models: a combined proof-of-mechanism and negative-results method study