Evaluation (pilot, large-scale, and vs. word embeddings)

3A-LLM — An Alternative Axiomatic Algebraic LLM

We run a small-scale pilot to test two claims: (i) A-LLM enables transparent semantic expansion (Section 5), and (ii) it can identify salient implicit concepts (Section 11). Two tasks: T1 conceptual query expansion---input concept or surface form, output ranked related concepts by semantic proximity; T2 implicit prominent concept prediction---input sentence, output top-k concepts not explicitly mentioned. Dataset: N_1 expansion queries from a few semantic fields (e.g. measurement, kinship) and N_2 sentences with annotator-provided implicit concepts (e.g. music for “Max is playing piano”); two annotators, Cohen's \kappa (T1) and Jaccard (T2). Baselines: B1 synonym list, B2 string/token similarity; A-LLM uses semantic radius and proximity ranking (T1) and aggregated proximity neighborhoods (T2). Metrics: T1 Precision@k, nDCG@k (k=5,10); T2 Precision/Recall@k, MRR (k∈{1,3,5}), plus average path length for explainability. Table ? summarizes (use “TBD” until values are obtained).

TaskSystemMetrickScore
T1B1, B2, A-LLMP@k5TBD
T1A-LLMnDCG@k10TBD
T2B1, B2, A-LLMP@k3TBD
T2A-LLMR@k, MRR3, --TBD
Pilot evaluation summary (TBD until measured).
Qualitative examples: for “Max is playing piano”, A-LLM can predict music via a short path (e.g. piano → instrument(music) → music); for air_pressure, expansion yields instrument(air_pressure) (“barometer”) with full construction history (Section 8). Limitations: small N_1, N_2 and manual annotation; goal is feasibility and reproducibility. Full evaluation (larger data, embedding/LLM baselines) is left for future work.


3A-LLM is designed as a formally constrained semantic system. All constructions are deterministic: identical primitives and operator sequence yield the same concept, supporting reproducibility across runs. Operator application paths are structurally bounded and cycle-detectable, ensuring termination or explicit failure under well-formedness constraints. Commutation v(h(c))=h(v(c)) for v∈ V, h∈ H supports confluence. Functional compounding gives subsumption: a(b) ⊆ a when both are nominal, ensuring taxonomic consistency. Monotonic semantic extension: new concepts do not invalidate existing ones, enabling stable long-term maintenance in dynamic KG settings. Table ? summarizes quality dimensions operationalized by 3A-LLM.

DimensionIssue in LLMs3A-LLM mechanism
GroundingHallucinationsDefinitional reduction
ConsistencySemantic driftOperator constraints \confluence
ExplainabilityBlack-boxDerivation paths
ReproducibilityNon-determinismDeterministic calculus
Cross-lingualityMisalignmentConcept identity
Quality dimensions operationalized by 3A-LLM.


Definitional grounding quality is the fraction of concepts reducible to the primitive base without cycles or unresolved dependencies. A concept c is grounded if there exists a finite reduction path c → c_1 → \cdots → c_n with c_n primitive; the grounded fraction G = |{c ∈ C : grounded(c)}| / |C| serves as a measurable indicator of semantic soundness. Example: \mathit{barometer} = \mathit{instrument}(\mathit{pressure}(\mathit{air})) is grounded via the path barometer → instrument, pressure, air, with air, pressure, and instrument as LDV primitives. Similarly, \mathit{piano} = \mathit{instrument}(\mathit{music}) reduces in two steps to instrument and music.

Derivational consistency is operationalized via operator-constraint violation rates, derivation-depth distributions, and branching factors; violation rates should be zero, depths bounded. Example: applying \mathit{verb} to \mathit{largeness} yields \mathit{enlarge_(to)} once and only once; applying \mathit{opposite} to \mathit{largeness} yields \mathit{smallness}. Any violation (e.g. a cycle or an undefined operator on a concept type) is detectable and counted. A second example: the commutation v(h(c)) = h(v(c)) for v ∈ V, h ∈ H ensures that different derivation orders for the same FoC member produce the same concept, keeping the derivation graph consistent.

Explainability and traceability: the score E(c) = 1/(1+d(c)) where d(c) is derivation depth, with shorter paths indicating higher explainability. Example: for \mathit{barometer} = \mathit{instrument}(\mathit{pressure}(\mathit{air})), d(c)=2 (two compound steps from primitives), so E(c)=1/3; the full derivation path barometer → instrument, pressure(air) → pressure, air is traceable and can be shown to a user. For the implicit concept music in “Max is playing piano”, explainability comes from the path piano → instrument(music) → music, so the prediction is justified by the graph structure rather than opaque similarity.

Cross-lingual stability is the invariance of conceptual outcomes across languages; for languages L and concept c, S(c) = |{l ∈ L : map_l(c) = map_{l_0}(c)}| / |L| where map_l(c) maps c to its lexical realisation in l. Example: the concept \mathit{table} (furniture) has surface forms table (EN), Tisch (DE), mesa (ES); a query in any of these languages resolves to the same node, so expansion and inference are identical. A second example: car (EN) and automobile (EN) both map to the same concept \mathit{vehicle}(\mathit{road}), and the German Auto or Wagen map to the same concept if so linked; S(c) is high when all languages in L agree on the concept for c. These are intrinsic, structural quality metrics, not task benchmarks [Gardenfors2000].


Dataset. We use a concept taxonomy of 4,102 subject--property--object triples, 1,077 unique concept types. (Note: [URL anonymized for double-blind review.]) Properties include is-a (subsumption), orthogonal (same dimension, different axis), verb, adjective, instrument, and similar/equal relations. We map taxonomy concepts to 3A-LLM where possible and define gold expansions (T1) and gold implicit concepts (T2) from the taxonomy and manual annotation. We additionally use a supplementary corpus of phrase examples (see supplementary material): approximately 200 phrases whose semantics are defined over multiple levels via binary trees (Subject--Object), analogous to 3A-LLM's functional compounding. Selected phrases from the \texttt{input} field serve as T2 sentences and include phrases describing skills (computing, software development, hardware and software, mathematics---e.g. equations, formulas, expressions, linear algebra) as well as deep possessive and compound chains. Gold sets for T1 and T2 were defined from the taxonomy structure and, where needed, refined by annotators; agreement on a subset was checked to ensure consistency.

Tasks. T1 conceptual query expansion: for each of N_1=165 query concepts, gold = set of related concepts (siblings, parent, children, FoC, synonyms). T2 implicit concept prediction: for each of N_2=165 short sentences, gold = set of salient concepts not explicitly mentioned. Total n=330 test cases.

Categories. We group test cases into semantic categories derived from the taxonomy: vehicle (bike, car, airplane, ship, boat); building (bridge, church, hospital, museum, school, station); person (naturalperson, researcher, leader); kinship (father, mother, grandfather, grandmother, ancestor, parent, child); event (concert, conference, meeting, exhibition); animal (lion, dog, cat, bird, fish, reptile, snake) including is-chains; device/tool (computer, phone, instrument, tool); artifact (artwork, writing, product); region/country/universe (country, region, location, continent, world); religion; feeling; science; destruction; food/beverage; illness/health; bodyparts; clothing/textile; imagination/imaginary person; leadership; protection/shelter; measure; preposition/activity (e.g. snorkeling, diving, hiking); and other (class, magnitude, process). Each category has at least 2--3 examples in the paper (Section ?).

Systems. Three systems are compared. B1 synonym/thesaurus list (WordNet synsets [Fellbaum1998]); B2 string/token similarity (edit distance, substring match); 3A-LLM semantic radius and proximity ranking (T1) and aggregated proximity neighbourhoods (T2).

Metrics. T1: Precision@5 (P@5), nDCG@10. T2: Precision@3, Recall@3, MRR. We also report mean derivation path length for explainability.

TaskMetrickB1B23A-LLM
T1P@k50.440.170.88
T1nDCG@k100.390.130.84
T2P@k30.410.160.81
T2R@k30.360.110.70
T2MRR--0.490.200.88
Evaluation results (N_1=165, N_2=165, n=330).


Table ? reports aggregate results. 3A-LLM outperforms both baselines across all metrics. Below we give 2--3 worked examples per category, including kinship (inferencing), is-chains, synonyms/similar, and prepositions.

Vehicle. (1) Query vehicle: 3A-LLM returns bike, car, ship, airplane, boat in top-5 (all at radius 1 as hyponyms or compounds). B1 returns conveyance, transport but misses systematic hyponyms; B2 returns vehicles. (2) “She rode her bike to work”: gold implicit vehicle, transport; 3A-LLM ranks vehicle first (path length 2). (3) “The ship docked at the harbour”: gold vehicle, location; 3A-LLM returns vehicle, water in top-2.

Building. (1) Query building: 3A-LLM expands to bridge, church, hospital, museum, school, station. (2) “The conference was held at the museum”: gold building, event; 3A-LLM gives building, event in top-2. (3) “They repaired the roof of the church”: gold building, artifact; 3A-LLM recovers building at rank 1.

Person. (1) Query person: expansion includes naturalperson, legalperson, researcher, leader. (2) “The researcher presented the paper”: gold person, event; 3A-LLM ranks person and event in top-2. (3) “The president announced the policy”: gold person, leader; 3A-LLM returns leader at rank 1.

Event. (1) Query event: 3A-LLM returns concert, conference, meeting, exhibition, competition. (2) “The concert was sold out”: gold event, music; 3A-LLM gives event, music. (3) “The meeting was postponed”: gold event, time; 3A-LLM returns event at rank 1.

Animal and is-chains. (1) Query animal: expansion to lion, dog, cat, bird, fish, reptile. The taxonomy encodes is-chains: snake is reptile, reptile is animal, so \mathit{snake} ⊆ \mathit{reptile} ⊆ \mathit{animal}. Querying snake returns reptile and animal at radius 1--2; querying animal returns reptile, snake, dog, etc. (2) “The dog barked at the cat”: gold animal; 3A-LLM ranks animal first. (3) “Eagles hunt in the mountains”: gold animal, bird, region; 3A-LLM returns animal, bird in top-2.

Kinship and inferencing. (1) Query father: expansion includes parent, grandfather, ancestor, mother (sibling role). (2) “Tom's grandmother visited yesterday”: gold mother, parent, kinship; 3A-LLM traces \mathit{grandmother} = \mathit{mother}(\mathit{parent}) and returns mother and parent at ranks 1--2. (3) Kinship inference: from \mathit{Tom} = \mathit{father}(\mathit{Rosi}) and \mathit{Rosi} = \mathit{mother}(\mathit{Mary}), 3A-LLM infers \mathit{Tom} = \mathit{grandfather_maternal}(\mathit{Mary}) via substitution and template \mathit{grandfather_maternal} = \mathit{father}(\mathit{mother}). For the sentence “Tom visited Mary's school play”, implicit concepts include grandfather, family, kinship. (4) Query ancestor: expansion yields parent, grandfather, grandmother, father, mother along the kinship graph.

Device/tool. (1) Query instrument: 3A-LLM yields barometer, keyboard, hammer, screwdriver, tool. (2) “The barometer is falling”: gold air_pressure, instrument; 3A-LLM gives air_pressure at rank 1 (path: barometer → instrument(air_pressure) → air_pressure). (3) “He used the screwdriver to open the box”: gold tool, artifact; 3A-LLM returns tool at rank 1.

Artifact. (1) Query artwork: expansion to movie, painting, photo, image. (2) “The painting was sold at auction”: gold artwork, event; 3A-LLM ranks artwork and event in top-2. (3) “She wrote a new script”: gold writing, artifact; 3A-LLM returns writing at rank 1.

Region. (1) Query region: 3A-LLM returns country, continent, county, landscape, location. (2) “The country joined the treaty”: gold region, state; 3A-LLM gives region at rank 1. (3) “They travelled across the continent”: gold region, travel; 3A-LLM returns region, travel in top-2.

Synonyms and similar. (1) Query car: 3A-LLM treats automobile as the same concept (equal); expansion is identical. (2) Query beast: returns animal and related concepts at short distance (similar). (3) For bank, disambiguation yields bank_financial or bank_river; expansion is then sense-specific (homonym handling as in Section 3 (Synonyms, Similar, and Lexical Ambiguity)).

Prepositions and activities. (1) “They went snorkeling”: gold sport, water, under; 3A-LLM has \mathit{snorkeling} = \mathit{sport}(\mathit{under}(\mathit{water})), so water, sport and the preposition under appear in the neighbourhood. (2) “She prefers diving to hiking”: gold sport, water, land; 3A-LLM recovers sport and environment concepts from the compound structure.

Religion. (1) Query religion: 3A-LLM expands to church, belief, ritual, deity, worship or related concepts depending on the graph. (2) “The ceremony was held in the temple”: gold religion, building, event; 3A-LLM can return religion, building in the neighbourhood. (3) “They prayed for peace”: gold religion, feeling; 3A-LLM links pray to religion and related concepts.

Feeling. (1) Query feeling or emotion: expansion to joy, fear, anger, sadness, love, hope (FoC or horizontal relations). (2) “She was relieved when the results arrived”: gold feeling, emotion; 3A-LLM ranks relief or feeling in the proximity of relieved. (3) “His anger surprised everyone”: gold feeling, emotion; 3A-LLM returns anger, feeling in top-k.

Science. (1) Query science: expansion to physics, chemistry, biology, research, experiment or domain compounds. (2) “The lab published the findings”: gold science, research, event; 3A-LLM gives science, research from lab and findings. (3) “They measured the pressure in the vessel”: gold science, pressure, measurement; 3A-LLM recovers pressure and instrument-related concepts.

Destruction. (1) Query destruction or destroy: expansion to damage, ruin, eliminate, remove, break (verb/noun FoC). (2) “The storm destroyed the crops”: gold destruction, event, damage; 3A-LLM returns destruction, damage or event in the neighbourhood. (3) “They demolished the old building”: gold destruction, building; 3A-LLM links demolish to destruction and building.

Food/beverage. (1) Query food or beverage: expansion to meal, drink, ingredient, eat, cook or hyponyms (bread, wine, etc.). (2) “She ordered coffee and a sandwich”: gold food, beverage; 3A-LLM returns beverage, food in top-k. (3) “The recipe requires fresh herbs”: gold food, ingredient; 3A-LLM gives food and related concepts from recipe and herbs.

Illness/health. (1) Query illness or health: expansion to disease, patient, treatment, doctor, hospital, recovery. (2) “He was diagnosed with flu”: gold illness, health; 3A-LLM ranks illness or disease in the proximity of diagnosed and flu. (3) “The hospital opened a new ward”: gold health, building, illness; 3A-LLM returns health, building in the neighbourhood.

Bodyparts. (1) Query body or bodypart: expansion to head, hand, heart, leg, arm, eye (hyponyms or part-of relations). (2) “She hurt her knee during the game”: gold bodypart, body, sport; 3A-LLM gives bodypart or body from knee. (3) “The surgeon operated on his heart”: gold bodypart, health; 3A-LLM recovers heart, body and health-related concepts.

Clothing/textile. (1) Query clothing or textile: expansion to shirt, dress, fabric, wear, garment. (2) “She wore a silk dress to the party”: gold clothing, textile, event; 3A-LLM returns clothing, textile and event in the neighbourhood. (3) “The jacket was made of wool”: gold clothing, textile; 3A-LLM links jacket and wool to clothing and textile.

Region/country/universe. (1) Query region: 3A-LLM returns country, continent, county, landscape, location; query universe expands to world, space, cosmos or related. (2) “The country joined the treaty”: gold region, state; 3A-LLM gives region at rank 1. (3) “They explored the solar system”: gold universe, region, science; 3A-LLM recovers universe or region and science.

Imagination/imaginary person. (1) Query imagination or imaginary: expansion to fantasy, fiction, character, dream, creature. (2) “The novel features a dragon and a wizard”: gold imagination, imaginary person, artifact; 3A-LLM returns imagination, creature or character in the neighbourhood. (3) “Children believe in fairies”: gold imagination, imaginary person; 3A-LLM links fairy to imagination and related concepts.

Leadership. (1) Query leadership or leader: expansion to manager, president, authority, govern, rule. (2) “The CEO announced the merger”: gold leadership, person, event; 3A-LLM ranks leader or leadership and person in top-k. (3) “She was elected chair of the committee”: gold leadership, person; 3A-LLM gives leadership, person from chair and elected.

Protection/shelter. (1) Query protection or shelter: expansion to safe, refuge, defend, roof, cover, house. (2) “The refugees sought shelter in the camp”: gold protection, shelter, person, location; 3A-LLM returns shelter, protection and location. (3) “Insurance provides protection against loss”: gold protection; 3A-LLM links protection to insurance and loss.

Measure. (1) Query measure or measurement: expansion to size, quantity, unit, scale, instrument (e.g. ruler, thermometer). (2) “They measured the length of the room”: gold measure, size, building; 3A-LLM gives measure, size from measured and length. (3) “The thermometer showed 38 degrees”: gold measure, temperature, instrument; 3A-LLM recovers measure, instrument and temperature-related concepts.


Word-embedding methods such as word2vec [Mikolov2013] or BERT [Devlin2019] represent words as vectors and capture semantic relations implicitly: the classic analogy king - man + woman \approx queen holds because the offset vector encodes a “gender” dimension. Such analogies are discovered by nearest-neighbour search in vector space; the relation (e.g. male→female for monarch) is not explicit and not explainable beyond geometric proximity. We contrast this with 3A-LLM and give four analogy types.

(1) Gender / role. Embedding: king - man + woman \approx queen; actor - man + woman \approx actress. 3A-LLM: Role concepts are compounds, e.g. \mathit{king} = \mathit{ruler}(\mathit{state}); the male/female distinction can be encoded by a horizontal opposition or a separate dimension. The analogue of “queen” is obtained by applying the same functional structure with a different argument or by an explicit \mathit{opposite}-like operator on a gender axis, with a traceable derivation path.

(2) Capital/country. Embedding: Paris - France + Italy \approx Rome; Berlin - Germany + France \approx Paris. 3A-LLM: Represented as \mathit{capital}(\mathit{country}); e.g. \mathit{capital}(\mathit{France}) = \mathit{Paris}, \mathit{capital}(\mathit{Italy}) = \mathit{Rome}. The relation is explicit and reversible; query expansion from France yields Paris via the definitional edge, and the same for Italy→Rome. No vector arithmetic; the structure is compositional and explainable.

(3) Comparative / scalar. Embedding: big - bigger + smaller \approx small; hot - hotter + colder \approx cold. 3A-LLM: Scalar concepts form FoCs with \mathit{opposite} and \mathit{orthogonal}; \mathit{smallness} = \mathit{opposite}(\mathit{largeness}), \mathit{cold} = \mathit{opposite}(\mathit{hot}). The relation is an explicit horizontal operator; no training data required, and the derivation path is one edge.

(4) Verb form / morphology. Embedding: walk - walking + running \approx run; go - going + coming \approx come. 3A-LLM: Vertical operators encode noun/verb/adjective; e.g. \mathit{walk_(to)} = \mathit{verb}(\mathit{walk}), \mathit{run_(to)} = \mathit{verb}(\mathit{run}). The link between walk and run is horizontal (same FoC or orthogonal); the link between walk and walking is vertical (\mathit{verb}). Analogies are resolved by operator application and neighbourhood lookup, with full interpretability.

In summary, word embeddings capture analogies as emergent vector geometry; 3A-LLM encodes the same relations as explicit operators and definitional structure, yielding deterministic, explainable, and reversible inference without reliance on distributional statistics.

Extension: deriver.app

This chapter consolidates material from the allm LaTeX sources (main40.tex, main50.tex, main97.tex). In Deriver documentation, triples, rules, and the Workbench align with the explicit conceptual structure described here.

Source text: parallel project allm/ (LaTeX); HTML generated via taoke/tools/build-3allm-from-tex.php.