Skip to content
Home

Lexeme (linguistics)

A lexeme is the abstract unit of lexical meaning in a language, grouping related wordforms under a single entry; it underlies dictionary headwords, morphological paradigms, and many NLP tasks.

Overview

A lexeme is an abstract unit of lexical meaning used in linguistics to group together related wordforms that share a common core sense. It is independent of particular inflected or derived forms: the various surface forms that a speaker produces (for example run, runs, ran, running) are treated as instances of one lexeme, conventionally represented by a citation form or lemma. The concept helps describe how vocabulary is organized in a language and how dictionaries and computational systems treat word families. For a general introduction to the linguistic concept see linguistics.

Core characteristics

Key properties that distinguish a lexeme include:

  • Abstractness: a lexeme denotes a set of related wordforms rather than any single surface form.
  • Independence from inflection: inflectional endings (tense, number, person) do not create new lexemes; they yield different wordforms of the same lexeme. See notes on inflectional morphology.
  • Multiword expressions: a lexeme can be a single word or a fixed multiword unit (for example, phrasal verbs like come in or idioms such as raining cats and dogs), which function as a single meaning-bearing unit in many analyses — compare discussion of single versus multiword entries at single or multiword.
  • Representative form (lemma): dictionaries typically choose one form to represent a lexeme, the lemma or citation form; see lemma.

History and terminology

The term lexeme arose in 20th-century structural and descriptive linguistics to give analysts a label for the ‘‘word’’ as an abstract entry in a language's mental lexicon. Etymologically it relates to the Greek root referring to speech or words, though modern usage is technical. Lexeme complements other terms such as morpheme (the minimal meaningful unit) and wordform or token (specific occurrences in speech or text).

Uses and applications

In lexicography and dictionaries, lexemes correspond to headwords: dictionary entries list a lemma and its inflected or derived wordforms and senses. Major English dictionaries include hundreds of thousands of lemmas and thereby cover many lexemes; see broad surveys at large dictionaries. In computational linguistics and natural language processing, lexemes underpin lemmatization (reducing surface forms to their base forms) and help with part-of-speech tagging, parsing, and building lexicons. For basic semantic description one may consult resources on lexical units and meaning at units of meaning.

Distinctions and notable facts

  1. Lexeme vs lemma: a lexeme is the abstract set; a lemma is the chosen citation form used to list that set in reference works. The lemma stands for the lexeme in dictionaries and corpora (lemma).
  2. Lexeme vs morpheme: morphemes are minimal units of form and meaning; a single lexeme may contain multiple morphemes (e.g. words with prefixes or suffixes) or be morphologically simple.
  3. When is a related form a different lexeme? Derived forms that produce distinct meanings (for example, teach vs teacher) are often treated as separate lexemes because the semantic relationship is not merely inflectional.
  4. Practical counts: counts of lexemes depend on criteria (inclusion of technical terms, multiword expressions, dialectal forms). Major English lexicons run into the hundreds of thousands, while the full set of lexical items in all registers is considerably larger.

Understanding lexemes clarifies how languages organize vocabulary at both theoretical and practical levels, informing dictionary compilation, annotation of corpora, and automated language processing systems.

Author

AlegsaOnline.com Lexeme (linguistics)

URL: https://en.alegsaonline.com/art/57620

Share