Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
📰 ArXiv cs.AI
arXiv:2601.16934v2 Announce Type: replace-cross Abstract: To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify any potential reflection biases, we introduce a permutation-based evaluation framework. With this, we observe that state-of-the-art embedding models exhibit systematic positional and language biases when documents are longer and consist of multiple segments. Specifically, early segments and segm
DeepCamp AI