Discussion

Controlled vocabulary of concepts (FCSS)

Discussion

  Word Embeddings:
A word embedding is an embedding in which words or other symbols are each assigned to a vector v with v ∈ ℝn. This is primarily used in machine learning. The goal here is to obtain an abstract representation of the meaning of the words or symbols while simultaneously reducing dimensions. Historically, one of the main limitations of static word embeddings or word vector space models is that words with multiple meanings are conflated into a single representation (a single vector in the semantic space). In other words, polysemy and homonymy are not handled properly. E.g., there is no possibility of distinguishing between homonyms, and it is important to note that synonyms must have the same value. So, our method offers a couple of features which we cannot get from word embeddings: We represent words with multiple meanings with different Word Sense Definitions. For homonyms we introduce different Basic Linguistic Symbols. The same holds for synonyms where each synonym can be related to another synonym with the Extra Concept Relation »SynonymOf. But the main difference is that word embeddings are not intended to compose the meaning of words from other words. That‘s why word embeddings cannot be used to implement Search by Meaning in the way we would need it.
  Natural Semantic Metalanguage (NSM):
Wierzbicka and Goddard in [GoWi2015a] state: The only 65 declared NSM primes have stabilized as a list of irreducible meanings, coded as English words with specific senses. These primes are hypothesized to be language universals, with most of them having been tested across a wide variety of languages without encountering disconfirmation. As we have shown in the Word Sense Definition (WSD) section for a concept c any number of WSDs are possible. Candidates for Semantic Primes have been identified by the authors in a decades lasting research on Minimal English (ME). In contrast to what the authors describe as Global English (GE) their Minimal English Lexicon (MEL) provides around 400 words on four layers: ME0 contains 65 semantic primes plus 100 variant forms thereof (allolexes). The layer ME1 has 70 universal or near-universal semantic compounds. The third layer ME2 adds 100 semantic compounds found in many languages, and ME3 collects 60 useful words for Minimal English which are not semantic compounds. According to the authors, one can assume that all translations of concepts using concepts of the layers ME0 and ME1, and partially even ME2 are universal or near-universal translatable to minimal lexicon versions of other languages like Minimal German, Minimal Finnish, Minimal Russian or Asian language like Minimal Chinese. This means that all language concepts that we have defined with WSDs based on concepts from the layers ME0 through ME2 can be automatically translated into all target languages without the need for lexicographer support. Compared to the translations of SynSets in WordNet this saves vast amounts of work.
  Longman´s Dictionary of Contemporary English:
All definitions in the dictionary are based on the Longman Defining Vocabulary (LDV) of around 2,000 common words. "The words in the LDV have been carefully chosen to ensure that the definitions are clear and easy to understand, and that the words used in explanations are easier than the words defined".  The following definitions for captain demonstrates, that a word can have several meanings (polysemy):  TDWcaptain = {the sailor in charge of a ship, the pilot in charge of an aircraft, a military officer with a fairly high rank, someone who leads a team or other group of people}. In our approach we have shown with the example of ωchurch_organ how to use Word Sense Definitons (WSD) such as ωship_captain, ωaircraft_pilot, ωmilitary_captain and ωteam_captain to avoid the disadvantages of polysemy. The words of the LDV have been used as the second layer of our Foundation Core Glossary (FCG).
  Princeton WordNet (PWN):
As we already discussed at the beginning, the Semantic Relations described in the methodology section that we have designated as Extra-Concept Relations (ECR) are essentially available in PWN. This also makes it possible to specify generic terms, synonyms, antonyms etc. However, in PWN there is no way of specifying a restriction in the sense of a Genus-Differentiae Pattern (GDP) for a term. Also, it is not possible to model concepts with the extended possibilities of Intra-Concept Relations (ICR) as discussed in the Concept Composition section. Consequently, in contrast to PWN, our Search by Meaning (SbM) methodology enables a Deep Semantic Search (DSS) that takes advantage of the composition of terms from others. For example, deep searching using queries like ‘air pressure’, ‘air gauge’, and ‘pressure gauge’ answers with the concept ‘barometer’, while PWN only finds the basic terms involved. On the other hand, DSS also returns the 'barometer' concept for all search terms combinations. PWN uses ILIs as unique identifiers for synsets, which are terms that can broadly be considered to represent synonyms. We, on the other hand, use speaking identifiers with the Basic Linguistic Symbols to make editing easier for lexicographers.
  Disambiguation of Prepositions and Possessives:
Using concept binary trees, the examples of [ScHw2018] can be modeled as follows: Using the absorption of the preposition ‘of’, Multiword prepositions, e.g., Δout_of = out, Δin_front_of = (in, front) are covered as well as idiomatic prepositional phrases like Δat_large = (at, large). The 50 super senses are organized in the three hierarchies circumstance (non-core properties of events), participant (entity playing a role in an event) and configuration (thing, usually an entity or property, involved in a static relationship to some other entity). Their Universal Semantic Tagset (UST) defines a cross-linguistic inventory of semantic classes for content and function words. Some more examples are Δafter_all = (after, all), Δas_well_as = ((as, well), as) and Prepositional Multi Word Expressions (PMWE) such as Δin_town =(in, town) and also non-prepositional MWE like Δtake_care = take_care.

Extension: deriver.app

Source: taoke.de — Discussion.