Nestor G Pestelos Jr · Reference · Print this page

Reference Guide

Embeddings

Published September 3, 2026 · Machine Learning & Representation Theory

Vector embeddings are dense numerical representations of discrete entities (such as tokens, sentences, code, or images) in a continuous high-dimensional vector space. By projecting discrete concepts into continuous geometries, embeddings enable machine learning models to quantify semantic similarity, capture associative analogies, and conduct sub-linear nearest neighbor retrieval across massive corpora.

1. Theoretical Motivation

In classical computational linguistics, words were treated as atomic symbols represented by orthogonal one-hot vectors [1]. In a vocabulary of size \(|V|\), each word \(w_i\) is represented by a vector \(\mathbf{x}_i \in \{0, 1\}^{|V|}\) where exactly one entry is 1 and all others are 0. This formulation exhibits two fundamental flaws:

Modern representation learning addresses this limitation through the distributional hypothesis, formulated by linguist J.R. Firth in 1957: "You shall know a word by the company it keeps" [2]. Dense vector embeddings map discrete vocabulary items into a low-dimensional continuous manifold \(\mathbb{R}^d\) (where typically \(d \in [256, 4096]\)), such that words appearing in similar contextual distributions are positioned close to one another in latent geometric space.

2. Mathematical Formulation & Geometry

Given an input token index \(i \in \{1, \dots, |V|\}\), the embedding operation corresponds to a matrix lookup in an embedding weight matrix \(\mathbf{W}_E \in \mathbb{R}^{|V| \times d}\):

$$\mathbf{e}_i = \mathbf{W}_E^T \mathbf{x}_i$$

where \(\mathbf{x}_i\) is the one-hot indicator vector for token \(i\), and \(\mathbf{e}_i \in \mathbb{R}^d\) is the resulting dense embedding.

Distance & Similarity Metrics

The semantic relationship between two embedded vectors \(\mathbf{u}, \mathbf{v} \in \mathbb{R}^d\) is quantified using geometric metrics in the embedding space:

Linear Substructure & Arithmetic

A seminal discovery in distributed representations is that continuous vector spaces encode analogical relationships as consistent directional offsets [3]. Relationships such as gender, grammatical tense, and capital-country associations emerge as stable translation vectors:

$$\mathbf{e}_{\text{king}} - \mathbf{e}_{\text{man}} + \mathbf{e}_{\text{woman}} \approx \mathbf{e}_{\text{queen}}$$

This linear property demonstrates that training objectives based on word co-occurrence implicitly perform matrix factorization on point-wise mutual information (PMI) matrices [4].

3. Architectural Evolution

Static Word Embeddings

Early distributed representation frameworks produced static lookup tables where each vocabulary token mapped to a single fixed vector regardless of context:

Contextual Sequence Embeddings

Static embeddings fail on polysemy: the token "bank" received the exact same vector in "river bank" and "investment bank." Modern deep architectures produce dynamically contextualized embeddings where each token's vector is a function of the entire sequence:

Positional Encodings

Because the self-attention mechanism is inherently permutation-invariant, positional information must be added directly into token representations [10]:

4. Geometric Pathologies: Anisotropy & Degeneration

A persistent challenge in deep language model representation is the anisotropy problem [12]. Rather than utilizing the full capacity of the \(d\)-dimensional hypersphere uniformly, learned token representations tend to collapse into a narrow, cone-shaped subspace. Consequently, arbitrary pairs of unrelated tokens often exhibit high positive cosine similarities (e.g. \(\cos > 0.7\)).

Remediations in modern embedding pipelines include:

5. Production Retrieval & Vector Indexing

In production applications such as Retrieval-Augmented Generation (RAG) and semantic search, systems query datasets containing millions or billions of embedding vectors. Exact \(k\)-nearest neighbor search via brute-force linear scanning requires \(O(N \cdot d)\) operations per query, which is computationally prohibitive at scale.

Production infrastructures deploy Approximate Nearest Neighbor (ANN) indexing structures:

See also

References

  1. [1] D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed. draft, Prentice Hall, 2024.
  2. [2] J. R. Firth, "A synopsis of linguistic theory 1930–1955," in Studies in Linguistic Analysis, Philological Society, Oxford, 1957, pp. 1–32.
  3. [3] T. Mikolov, K. Chen, G. Corrado, and J. Dean, "Efficient Estimation of Word Representations in Vector Space," in Proceedings of the International Conference on Learning Representations (ICLR), 2013. https://arxiv.org/abs/1301.3781
  4. [4] O. Levy and Y. Goldberg, "Neural Word Embedding as Implicit Matrix Factorization," in Advances in Neural Information Processing Systems (NeurIPS), vol. 27, 2014, pp. 2177–2185.
  5. [5] J. Pennington, R. Socher, and C. D. Manning, "GloVe: Global Vectors for Word Representation," in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1532–1543. https://nlp.stanford.edu/pubs/glove.pdf
  6. [6] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, "Enriching Word Vectors with Subword Information," Transactions of the Association for Computational Linguistics, vol. 5, pp. 135–146, 2017. https://arxiv.org/abs/1607.04606
  7. [7] M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, "Deep Contextualized Word Representations," in NAACL-HLT, 2018, pp. 2227–2237. https://arxiv.org/abs/1802.05365
  8. [8] J. Devlin, M. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in NAACL-HLT, 2019, pp. 4171–4186. https://arxiv.org/abs/1810.04805
  9. [9] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in EMNLP-IJCNLP, 2019, pp. 3982–3992. https://arxiv.org/abs/1908.10084
  10. [10] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention Is All You Need," in NeurIPS, 2017, pp. 5998–6008. https://arxiv.org/abs/1706.03762
  11. [11] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, "RoFormer: Enhanced Transformer with Rotary Position Embedding," Neurocomputing, vol. 568, p. 127063, 2024. https://arxiv.org/abs/2104.09864
  12. [12] B. Gao, Y. Song, S. Shen, et al., "Representation Degeneration Problem in Language Modeling," in ICLR, 2019. https://arxiv.org/abs/1907.12009
  13. [13] Y. A. Malkov and D. A. Yashunin, "Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, pp. 824–836, 2018. https://arxiv.org/abs/1603.09320