Evaluating Embeddings with MTEB Massive text embeddings benchmark - Nils Reimers

Cohere · Beginner ·📄 Research Papers Explained ·3y ago

Key Takeaways

The video discusses the evaluation of embeddings using the MTEB benchmark, introduced by Nils Reimers, and explores the applications of embedding models, including few-shot learning and prompt engineering with Sentence Transformers and GPT-3.

Full Transcript

this year we published the massive text embedding benchmarks um which is kind of like a follow-up work bit more you know more put it bit more broadly in perspective so we collected different tasks like where you can use embedding models you can use them for clustering you can use them for BX mining meaning finding sentence with the same meaning in different languages you can use them for retrieve for semantic textual similarity for summarization for text classification for pair classification and for reranking and so we so so this is like a project that started also a really long time ago so when I published the sentence bird paper sentence Informer papers it showed good results on STS and S eval which was like the defact standard in embedding evaluation but if you really use it didn't perform that that well and also the original sentence SP model didn't perform that well as universal sentence encoder even such that The Benchmark on STS showed the opposite and so over the years with different help of different people and so on uh I collected first internally like a lot more training data a lot more data sets to evaluate my models on then luckily this year we have been able to open Source this and run a lot of evaluation on this and here we tested different models on these 58 data sets uh what you see on the x-axis is the speed like how quick is the model on the y- AIS how good is the model and then the size of the circles like how big how how many output Dimensions do you have and yeah what you totally see is like some cluster around here here where people used like bird based like 110 billion parameter models then we see like different sizes like the one model on mpet and GTR they perform extremely well also here some minm models work really well and then at the more left side we have like these massive six billion parameter model that show um also kind of like a nice Improvement but if you really compare the difference between like this model so here the center point and the npet center point the difference is not so small um but it's like two two magnitudes slower than these smaller models so I always found like this interesting to find like a model that's not only good but also would it fast so another interesting application we showed this year is fot learning with embedding models so gpt3 popularize fre short learning with prompt and in Contex learning so so what opening I showed in their paper is that if you want to do sentiment classification you can create like a prompt like this describe the sentiment then you have some in context examples I love this place which you say it's positive I don't I don't go here which is negative then you have a new review great pizza which you want to classify and I was never a fan of prompting and in context learning because first it has an extremely high sensitivity to The Prompt so here it can make a difference if you put a colon or an exclamation mark or question mark which can absolutely change the the performance and so you never know do I have to put like a colon or a question mark at the end of this do I have to add like a new line or not um also writing the prompt can be challenging for sentiment classification it's kind of easy but if you have like more complex nuanced um tasks it can be challenging to to write the prompt for that further challenge is you have limited context lengths so the original gpt3 can do like up to like 2,000 word pieces so if you have like longer examples like movie reviews you might can only present like 10 training examples but what happens when you have like 20 training examples examples or 100 training examples there's like no way how you can use these in a in context learning example and also you have quite High compute overhead so the the encoding um time complexity is quadratic with the length uh of the input so if you have to provide like all these examples at every call you like a lot of compute overhead to it

Original Description

Sentence Transformers and Embedding Evaluation - Talking Language AI Ep#3 Full episode: https://youtu.be/apuDeylm1uE About The Speaker: Nils is the creator of Sentence-BERT and has authored several well-known research papers, including Sentence-BERT and the popular Sentence Transformers library. He’s also worked as a Research Scientist at HuggingFace, (co-)founded several web companies, and worked as an AI consultant in the area of investment banking, media, and IoT. === In our conversation, Nils gives us an introduction to the Sentence-BERT package and the large language models provided in it. He also shares some lessons from his experience in open-source development of such a popular package. Finally, Nils touches on his research collaborations on how to evaluate embeddings through works like MTEB: Massive Text Embedding Benchmark and BEIR. To go deeper into these tools, and other concepts around embeddings, watch the video and join the conversation on Discord. Stay tuned for more episodes in our Talking Language AI series! === Join the Cohere Discord: https://discord.gg/co-mmunity Discussion thread for this episode (feel free to ask questions): https://discord.com/channels/954421988141711382/1052547510910062624 Watch more episodes of Talking Language AI: https://www.youtube.com/playlist?list=PLLalUvky4CLJ9ZgtZguDJ7dAYuI1bfaYW === Resources: Bonjour. مرحبا. Guten tag. Hola. Cohere's Multilingual Text Understanding Model is Now Available: https://txt.cohere.ai/multilingual/ SBERT: https://www.sbert.net/ SBERT Paper: https://arxiv.org/abs/1908.10084 MTEB: Massive Text Embedding Benchmark: https://arxiv.org/abs/2210.07316 BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models: https://openreview.net/forum?id=wCu6T5xFjeJ SetFit - Efficient Few-shot Learning with Sentence Transformers https://github.com/huggingface/setfit
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Cohere · Cohere · 50 of 60

1 Andreas Madsen on Independent Research and Interpretability
Andreas Madsen on Independent Research and Interpretability
Cohere
2 Plex: Towards Reliability using Pretrained Large Model Extensions
Plex: Towards Reliability using Pretrained Large Model Extensions
Cohere
3 Independent Research Panel Discussion
Independent Research Panel Discussion
Cohere
4 The Future of ML Ops: Open Challenges and Opportunities
The Future of ML Ops: Open Challenges and Opportunities
Cohere
5 C4AI Special - Grad School Applications
C4AI Special - Grad School Applications
Cohere
6 Cohere For AI Fireside Chat: Samy Bengio
Cohere For AI Fireside Chat: Samy Bengio
Cohere
7 Cohere For AI - Scholars Program Information Session
Cohere For AI - Scholars Program Information Session
Cohere
8 Modular and Composable Transfer Learning with Jonas Pfeiffer
Modular and Composable Transfer Learning with Jonas Pfeiffer
Cohere
9 Jay Alammar Presents Large Language Models for Real World Applications
Jay Alammar Presents Large Language Models for Real World Applications
Cohere
10 Catherine Olsson - Mechanistic Interpretability: Getting Started
Catherine Olsson - Mechanistic Interpretability: Getting Started
Cohere
11 How To Prompt Engineer a Tech Interview App | TOHacks 2022 Winners
How To Prompt Engineer a Tech Interview App | TOHacks 2022 Winners
Cohere
12 C4AI Sparks: Samy Bengio
C4AI Sparks: Samy Bengio
Cohere
13 BERTopic for Topic Modeling - Maarten Grootendorst - Talking Language AI Ep#1
BERTopic for Topic Modeling - Maarten Grootendorst - Talking Language AI Ep#1
Cohere
14 Exploring News Headlines With Text Clustering | Jay Alammar
Exploring News Headlines With Text Clustering | Jay Alammar
Cohere
15 Scale TransformX | Fireside Chat: Aidan Gomez and Alexandr Wang
Scale TransformX | Fireside Chat: Aidan Gomez and Alexandr Wang
Cohere
16 Making Large Language Models Accessible | Scale AI Fireside chat with Bill MacCartney
Making Large Language Models Accessible | Scale AI Fireside chat with Bill MacCartney
Cohere
17 Intro to KeyBERT - BERTopic for Topic Modeling
Intro to KeyBERT - BERTopic for Topic Modeling
Cohere
18 Intro to PolyFuzz - BERTopic for Topic Modeling
Intro to PolyFuzz - BERTopic for Topic Modeling
Cohere
19 API Design Philosophy - BERTopic for Topic Modeling
API Design Philosophy - BERTopic for Topic Modeling
Cohere
20 Code demo of BERTopic - BERTopic for Topic Modeling
Code demo of BERTopic - BERTopic for Topic Modeling
Cohere
21 Short texts vs long texts in BERTopic- BERTopic for Topic Modeling
Short texts vs long texts in BERTopic- BERTopic for Topic Modeling
Cohere
22 How People can help BERTopic - BERTopic for Topic Modeling
How People can help BERTopic - BERTopic for Topic Modeling
Cohere
23 Cohere For AI: Training Sensorimotor Agency in Cellular Automata with Bert Chan
Cohere For AI: Training Sensorimotor Agency in Cellular Automata with Bert Chan
Cohere
24 Cohere API Community Demos | October 2022
Cohere API Community Demos | October 2022
Cohere
25 Perfect Prompt Demo By Arjun Patel
Perfect Prompt Demo By Arjun Patel
Cohere
26 Project Idea Generator Demo By Tobechukwu Okamkpa
Project Idea Generator Demo By Tobechukwu Okamkpa
Cohere
27 SuperTransformer Demo By Amir Nagri and Team Megatron
SuperTransformer Demo By Amir Nagri and Team Megatron
Cohere
28 Cohere For AI Fireside Chat: Pablo Samuel Castro
Cohere For AI Fireside Chat: Pablo Samuel Castro
Cohere
29 How Startups Can Use NLP to Build a Competitive Moat
How Startups Can Use NLP to Build a Competitive Moat
Cohere
30 Build Chatbots Faster with Large Language Models
Build Chatbots Faster with Large Language Models
Cohere
31 Tools to Improve Training Data - Vincent Warmerdam - Talking Language AI Ep#2
Tools to Improve Training Data - Vincent Warmerdam - Talking Language AI Ep#2
Cohere
32 Utku Evci - Sparsity and Beyond Static Network Architectures
Utku Evci - Sparsity and Beyond Static Network Architectures
Cohere
33 Adding human intelligence to ML models with human-learn #shorts #machinelearning #nlp
Adding human intelligence to ML models with human-learn #shorts #machinelearning #nlp
Cohere
34 Iterating on your data with doubtlab - Tools to Improve Training Data
Iterating on your data with doubtlab - Tools to Improve Training Data
Cohere
35 Adding Human Intelligence to ML models with Human learn - Tools to Improve Training Data
Adding Human Intelligence to ML models with Human learn - Tools to Improve Training Data
Cohere
36 Scikt Learn embeddings helpers with Embetter - Tools to Improve Training Data
Scikt Learn embeddings helpers with Embetter - Tools to Improve Training Data
Cohere
37 Building Cohere API Demo App With Streamlit | Adrien Morisot
Building Cohere API Demo App With Streamlit | Adrien Morisot
Cohere
38 Rosanne Liu - career creation for non-standard candidates
Rosanne Liu - career creation for non-standard candidates
Cohere
39 Giving computers many human languages with Cohere's multilingual embeddings
Giving computers many human languages with Cohere's multilingual embeddings
Cohere
40 Learning by Distilling Context with Charlie Snell
Learning by Distilling Context with Charlie Snell
Cohere
41 Sentence Transformers and Embedding Evaluation - Nils Reimers - Talking Language AI Ep#3
Sentence Transformers and Embedding Evaluation - Nils Reimers - Talking Language AI Ep#3
Cohere
42 Reflecting on for.ai...
Reflecting on for.ai...
Cohere
43 Create a Custom Language Model with Surge AI and Cohere
Create a Custom Language Model with Surge AI and Cohere
Cohere
44 Cohere API Community Demos | November 2022
Cohere API Community Demos | November 2022
Cohere
45 Cohere API Community Demos | December 2022
Cohere API Community Demos | December 2022
Cohere
46 Cohere For AI Presents: Colin Raffel
Cohere For AI Presents: Colin Raffel
Cohere
47 Lucas Beyer - FlexiViT: One Model for All Patch Sizes
Lucas Beyer - FlexiViT: One Model for All Patch Sizes
Cohere
48 What is Neural Search? Nils Reimers - Sentence Transformers and Embedding Evaluation
What is Neural Search? Nils Reimers - Sentence Transformers and Embedding Evaluation
Cohere
49 Evaluating Information Retrieval with BEIR
Evaluating Information Retrieval with BEIR
Cohere
Evaluating Embeddings with MTEB Massive text embeddings benchmark - Nils Reimers
Evaluating Embeddings with MTEB Massive text embeddings benchmark - Nils Reimers
Cohere
51 High quality text classification with few training examples with SetFit
High quality text classification with few training examples with SetFit
Cohere
52 Multilingual and cross lingual embeddings - Nils Reimers
Multilingual and cross lingual embeddings - Nils Reimers
Cohere
53 Developing open-source software: lessons, benefits, and challenges - Nils Reimers
Developing open-source software: lessons, benefits, and challenges - Nils Reimers
Cohere
54 Ask Me Anything with Ed Grefenstette, Head of Machine Learning at Cohere
Ask Me Anything with Ed Grefenstette, Head of Machine Learning at Cohere
Cohere
55 HyperWrite Powers Its Generative AI Service with Cohere
HyperWrite Powers Its Generative AI Service with Cohere
Cohere
56 EMNLP 2022 Conference Special Edition - Talking Language AI #4
EMNLP 2022 Conference Special Edition - Talking Language AI #4
Cohere
57 Cohere API Community Demos | January 2023
Cohere API Community Demos | January 2023
Cohere
58 C4AI Sparks: Rosanne Liu on Career Creation for Non-Standard Candidates
C4AI Sparks: Rosanne Liu on Career Creation for Non-Standard Candidates
Cohere
59 Michael Tschannen -  Image-and-Language Understanding from Pixels Only
Michael Tschannen - Image-and-Language Understanding from Pixels Only
Cohere
60 How to Add AI to your App
How to Add AI to your App
Cohere

The video teaches how to evaluate embeddings using the MTEB benchmark and explores the applications of embedding models, including few-shot learning and prompt engineering with Sentence Transformers and GPT-3. The speaker, Nils Reimers, discusses the challenges of prompting and in-context learning, and presents alternative approaches using embedding models.

Key Takeaways
  1. Collect different tasks for embedding model evaluation
  2. Use MTEB benchmark to evaluate embedding models
  3. Apply Sentence Transformers for embedding evaluation
  4. Use GPT-3 for few-shot learning
  5. Explore alternative approaches to prompting and in-context learning
💡 Embedding models can be effectively evaluated using the MTEB benchmark, and can be applied to various tasks, including few-shot learning and prompt engineering, with alternative approaches to prompting and in-context learning.

Related Reads

📰
6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation
Only 6.5% of neuro-symbolic AI literature can be reproduced from published artifacts, highlighting the need for a reproducibility audit framework
ArXiv cs.AI
📰
Research Publications, Patents & Innovation Output at Quantum University
Boost academic reputation through research output and innovation at Quantum University
Medium · Machine Learning
📰
Every Researcher Should Start Managing Research Intelligence Assets™
Researchers can generate long-term value from their intellectual assets by managing them effectively, which is crucial for maximizing research impact
Medium · AI
📰
AC comment and our reply disappeared on OpenReview [D]
Learn how to troubleshoot missing comments on OpenReview and understand the potential causes of disappeared posts
Reddit r/MachineLearning
Up next
Why the Best Ideas Can't Be Neatly Explained
David Perell
Watch →