diff --git a/tutorials/README.md b/tutorials/README.md deleted file mode 100644 index a1dd0533..00000000 --- a/tutorials/README.md +++ /dev/null @@ -1,16 +0,0 @@ -
- - Tutorials - - -
- -# Tutorials - -This is a list of all our tutorials. They are all self-contained ipython notebooks. - -| | what? | Link | -|--------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------| -| **Recipe search** 🍝 | Learn how to do lightning-fast semantic search by distilling a small model. Compare a really tiny model to a larger with one with a better vocabulary. Learn what Fattoush is (delicious). | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/minishlab/model2vec/blob/master/tutorials/recipe_search.ipynb) | -| **Semantic chunking** đŸ§© | Learn how to chunk your text into meaningful segments with [Chonkie](https://github.com/chonkie-inc/chonkie) at lightning-speed. Efficiently query your chunks with [Vicinity](https://github.com/MinishLab/vicinity). | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/minishlab/model2vec/blob/master/tutorials/semantic_chunking.ipynb) | -| **Training a classifier** đŸ§© | Learn how to train a classifier using model2vec. Lightning fast, great performance, especially on small datasets | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/minishlab/model2vec/blob/master/tutorials/train_classifier.ipynb) | diff --git a/tutorials/recipe_search.ipynb b/tutorials/recipe_search.ipynb deleted file mode 100644 index 6158d5ce..00000000 --- a/tutorials/recipe_search.ipynb +++ /dev/null @@ -1,500 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Recipe Search using Model2Vec**\n", - "\n", - "This notebook demonstrates how to use the Model2Vec library to search for recipes based on a given query. We will use the [recipe dataset](https://huggingface.co/datasets/Shengtao/recipe).\n", - "We will be using the `model2vec` in different modes to search for recipes based on a query, using both our own pre-trained models, as well as a domain-specific model we will distill ourselves in this tutorial.\n", - "\n", - "Three modes of Model2Vec use are demonstrated:\n", - "1. **Using a pre-trained output vocab model**: Uses a pre-trained output embedding model. This is a very small model that uses a subword tokenizer. \n", - "2. **Using a pre-trained glove vocab model**: Uses pre-trained glove vocab model. This is a larger model that uses a word tokenizer.\n", - "3. **Using a custom vocab model**: Uses a custom domain-specific vocab model that is distilled on a vocab created from the recipe dataset. " - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "# Install the necessary libraries\n", - "!pip install numpy datasets scikit-learn transformers model2vec\n", - " \n", - "# Import the necessary libraries\n", - "import regex\n", - "from collections import Counter\n", - "\n", - "import numpy as np\n", - "from datasets import load_dataset\n", - "from sklearn.metrics import pairwise_distances\n", - "from tokenizers.pre_tokenizers import Whitespace\n", - "\n", - "from model2vec import StaticModel\n", - "from model2vec.distill import distill" - ] - }, - { - "cell_type": "code", - "execution_count": 96, - "metadata": {}, - "outputs": [], - "source": [ - "# Load the recipe dataset\n", - "dataset = load_dataset(\"Shengtao/recipe\", split=\"train\")\n", - "# Convert the dataset to a pandas DataFrame\n", - "dataset = dataset.to_pandas()\n", - "# Take the title column as our recipes corpus\n", - "recipes = dataset[\"title\"]" - ] - }, - { - "cell_type": "code", - "execution_count": 97, - "metadata": {}, - "outputs": [ - { - "data": { - "text/html": [ - "
\n", - "\n", - "\n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - " \n", - "
titlecategorydescriptioningredientsdirections
0Simple Macaroni and Cheesemain-dishA very quick and easy fix to a tasty side-dish...1 (8 ounce) box elbow macaroni ; Œ cup butter ...Bring a large pot of lightly salted water to a...
1Gourmet Mushroom Risottomain-dishAuthentic Italian-style risotto cooked the slo...6 cups chicken broth, divided ; 3 tablespoons ...In a saucepan, warm the broth over low heat. W...
2Dessert Crepesbreakfast-and-brunchEssential crepe recipe. Sprinkle warm crepes ...4 eggs, lightly beaten ; 1 ⅓ cups milk ; 2 ta...In large bowl, whisk together eggs, milk, melt...
3Pork Steaksmeat-and-poultryMy mom came up with this recipe when I was a c...Œ cup butter ; Œ cup soy sauce ; 1 bunch green...Melt butter in a skillet, and mix in the soy s...
4Quick and Easy Pizza CrustbreadThis is a great recipe when you don't want to ...1 (.25 ounce) package active dry yeast ; 1 tea...Preheat oven to 450 degrees F (230 degrees C)....
\n", - "
" - ], - "text/plain": [ - " title category \\\n", - "0 Simple Macaroni and Cheese main-dish \n", - "1 Gourmet Mushroom Risotto main-dish \n", - "2 Dessert Crepes breakfast-and-brunch \n", - "3 Pork Steaks meat-and-poultry \n", - "4 Quick and Easy Pizza Crust bread \n", - "\n", - " description \\\n", - "0 A very quick and easy fix to a tasty side-dish... \n", - "1 Authentic Italian-style risotto cooked the slo... \n", - "2 Essential crepe recipe. Sprinkle warm crepes ... \n", - "3 My mom came up with this recipe when I was a c... \n", - "4 This is a great recipe when you don't want to ... \n", - "\n", - " ingredients \\\n", - "0 1 (8 ounce) box elbow macaroni ; ÂŒ cup butter ... \n", - "1 6 cups chicken broth, divided ; 3 tablespoons ... \n", - "2 4 eggs, lightly beaten ; 1 ⅓ cups milk ; 2 ta... \n", - "3 ÂŒ cup butter ; ÂŒ cup soy sauce ; 1 bunch green... \n", - "4 1 (.25 ounce) package active dry yeast ; 1 tea... \n", - "\n", - " directions \n", - "0 Bring a large pot of lightly salted water to a... \n", - "1 In a saucepan, warm the broth over low heat. W... \n", - "2 In large bowl, whisk together eggs, milk, melt... \n", - "3 Melt butter in a skillet, and mix in the soy s... \n", - "4 Preheat oven to 450 degrees F (230 degrees C).... " - ] - }, - "execution_count": 97, - "metadata": {}, - "output_type": "execute_result" - } - ], - "source": [ - "# Display the first few rows of the dataset for the specified columns\n", - "dataset[[\"title\", \"category\", \"description\", \"ingredients\", \"directions\"]].head()" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "First, we will set up a function to handle similarity search that we can use in this tutorial." - ] - }, - { - "cell_type": "code", - "execution_count": 46, - "metadata": {}, - "outputs": [], - "source": [ - "# Define a function to find the most similar titles in a dataset to a given query\n", - "def find_most_similar_items(model: StaticModel, embeddings: np.ndarray, query: str, top_k=5) -> list[tuple[int, float]]:\n", - " \"\"\"\n", - " Finds the most similar items in a dataset to the given query using the specified model.\n", - "\n", - " :param model: The model used to generate embeddings.\n", - " :param embeddings: The embeddings of the dataset.\n", - " :param query: The query recipe title.\n", - " :param top_k: The number of most similar titles to return.\n", - " :return: A list of tuples containing the most similar titles and their cosine similarity scores.\n", - " \"\"\"\n", - " # Generate embedding for the query\n", - " query_embedding = model.encode(query)[None, :]\n", - "\n", - " # Calculate pairwise cosine distances between the query and the precomputed embeddings\n", - " distances = pairwise_distances(query_embedding, embeddings, metric='cosine')[0]\n", - "\n", - " # Get the indices of the most similar items (sorted in ascending order because smaller distances are better)\n", - " most_similar_indices = np.argsort(distances)\n", - "\n", - " # Convert distances to similarity scores (cosine similarity = 1 - cosine distance)\n", - " most_similar_scores = [1 - distances[i] for i in most_similar_indices[:top_k]]\n", - "\n", - " # Return the top-k most similar indices and similarity scores\n", - " return list(zip(most_similar_indices[:top_k], most_similar_scores))" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Using a pre-trained output vocab model**\n", - "\n", - "In this part, we will use a pre-trained output vocab model to encode the recipes and search using multiple queries. The output vocab model is very small and fast while still providing good results. Since the model uses a sub-word tokenizer, it is able to handle out-of-vocabulary words and provide good results even for words that are not in the base vocab." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "# Load the M2V output model from the HuggingFace hub\n", - "model_name = \"minishlab/M2V_base_output\"\n", - "model_output = StaticModel.from_pretrained(model_name)" - ] - }, - { - "cell_type": "code", - "execution_count": 91, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Most similar recipes to 'cheeseburger':\n", - "Title: `Double Cheeseburger`, Similarity Score: 0.9028\n", - "Title: `Cheeseburger Chowder`, Similarity Score: 0.8574\n", - "Title: `Cheeseburger Sliders`, Similarity Score: 0.8413\n", - "Title: `Cheeseburger Salad`, Similarity Score: 0.8384\n", - "Title: `Cheeseburger Soup I`, Similarity Score: 0.8298\n", - "\n", - "Most similar recipes to 'fattoush':\n", - "Title: `Fattoush`, Similarity Score: 1.0000\n", - "Title: `Lebanese Fattoush`, Similarity Score: 0.8370\n", - "Title: `Aunty Terese's Fattoush`, Similarity Score: 0.7630\n", - "Title: `Arabic Fattoush Salad`, Similarity Score: 0.7588\n", - "Title: `Authentic Lebanese Fattoush`, Similarity Score: 0.7584\n" - ] - } - ], - "source": [ - "# Find recipes using the output embeddings model\n", - "top_k = 5\n", - "\n", - "# Find the most similar recipes to the given queries\n", - "query = \"cheeseburger\"\n", - "embeddings = model_output.encode(recipes)\n", - "\n", - "results = find_most_similar_items(model_output, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " \n", - "print()\n", - "\n", - "query = \"fattoush\"\n", - "results = find_most_similar_items(model_output, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " " - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "As can be seen, we get some good results for the queries. The model is able to find recipes that are similar to the query." - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Using a pre-trained output vocab model**\n", - "\n", - "In this part, we will use a pre-trained glove vocab model to encode the recipes and search using multiple queries. The glove vocab model is a bit larger and slower than the output vocab model but can provide better results. However, as we will see, it suffers from the out-of-vocabulary problem, since the glove vocab is not designed for the cooking recipe domain." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "# Load the M2V glove model from the HuggingFace hub\n", - "model_name = \"minishlab/M2V_base_glove\"\n", - "model_glove = StaticModel.from_pretrained(model_name)" - ] - }, - { - "cell_type": "code", - "execution_count": 92, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Most similar recipes to 'cheeseburger':\n", - "Title: `Double Cheeseburger`, Similarity Score: 0.8744\n", - "Title: `Cheeseburger Meatloaf`, Similarity Score: 0.8246\n", - "Title: `Cheeseburger Salad`, Similarity Score: 0.8160\n", - "Title: `Hearty American Cheeseburger`, Similarity Score: 0.8006\n", - "Title: `Cheeseburger Chowder`, Similarity Score: 0.7989\n", - "\n", - "Most similar recipes to 'fattoush':\n", - "Title: `Simple Macaroni and Cheese`, Similarity Score: 0.0000\n", - "Title: `Fresh Tomato and Cucumber Salad with Buttery Garlic Croutons`, Similarity Score: 0.0000\n", - "Title: `Grilled Cheese, Apple, and Thyme Sandwich`, Similarity Score: 0.0000\n", - "Title: `Poppin' Turkey Salad`, Similarity Score: 0.0000\n", - "Title: `Chili - The Heat is On!`, Similarity Score: 0.0000\n" - ] - } - ], - "source": [ - "# Find recipes using the output embeddings model\n", - "top_k = 5\n", - "\n", - "# Find the most similar recipes to the given queries\n", - "query = \"cheeseburger\"\n", - "embeddings = model_glove.encode(recipes)\n", - "\n", - "results = find_most_similar_items(model_glove, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " \n", - "print()\n", - "\n", - "query = \"fattoush\"\n", - "results = find_most_similar_items(model_glove, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " " - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "As can be seen, we get good results when we search for an in vocab query (`cheeseburger`), but when we search for an out-of-vocab query (`fattoush`), the model is not able to find any relevant recipes. To fix this, we will now distill a custom vocab model on the recipe dataset." - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Using a custom vocab model**\n", - "\n", - "In this part, we will distill a custom vocab model on the recipe dataset and use it to encode the recipes and search using multiple queries. This will create a domain-specific model2vec model. First, we will set up a function to create a vocabulary from a list of texts (in our case, a list of recipe titles)." - ] - }, - { - "cell_type": "code", - "execution_count": 98, - "metadata": {}, - "outputs": [], - "source": [ - "# Set up a regex tokenizer to split texts into words and punctuation\n", - "my_regex = regex.compile(r\"\\w+|[^\\w\\s]+\")\n", - "\n", - "def create_vocab(texts: list[str], tokenizer: Whitespace, size: int = 30_000) -> list[str]:\n", - " \"\"\"\n", - " Create a vocab from a list of texts.\n", - " \n", - " :param texts: A list of texts.\n", - " :param tokenizer: A whitespace tokenizer.\n", - " :param size: The size of the vocab.\n", - " :return: A vocab sorted by frequency.\n", - " \"\"\"\n", - " counts = Counter()\n", - " for text in texts:\n", - " tokens = tokenizer.pre_tokenize_str(text.lower())\n", - " tokens = [token for token, _ in tokens]\n", - " counts.update(tokens)\n", - " vocab = [word for word, _ in counts.most_common(size)]\n", - " return vocab" - ] - }, - { - "cell_type": "code", - "execution_count": 88, - "metadata": {}, - "outputs": [ - { - "name": "stderr", - "output_type": "stream", - "text": [ - "100%|██████████| 8/8 [00:08<00:00, 1.04s/it]\n" - ] - } - ], - "source": [ - "# Choose a Sentence Transformer model and a tokenizer\n", - "model_name = \"BAAI/bge-small-en-v1.5\"\n", - "tokenizer = Whitespace()\n", - "\n", - "# Create a custom vocab from the recipe titles\n", - "vocab = create_vocab(recipes, tokenizer)\n", - "\n", - "# Distill a model2vec model using the Sentence Transformer model and the custom vocab\n", - "model_custom = distill(model_name=model_name, vocabulary=vocab, pca_dims=256)" - ] - }, - { - "cell_type": "code", - "execution_count": 93, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Most similar recipes to 'cheeseburger':\n", - "Title: `Cheeseburger Salad`, Similarity Score: 0.9528\n", - "Title: `Cheeseburger Casserole`, Similarity Score: 0.9030\n", - "Title: `Cheeseburger Chowder`, Similarity Score: 0.8635\n", - "Title: `Cheeseburger Pie`, Similarity Score: 0.8401\n", - "Title: `Cheeseburger Meatloaf`, Similarity Score: 0.8184\n", - "\n", - "Most similar recipes to 'fattoush':\n", - "Title: `Fattoush`, Similarity Score: 1.0000\n", - "Title: `Fatoosh`, Similarity Score: 0.7488\n", - "Title: `Lebanese Fattoush`, Similarity Score: 0.6344\n", - "Title: `Arabic Fattoush Salad`, Similarity Score: 0.6108\n", - "Title: `Fattoush (Lebanese Salad)`, Similarity Score: 0.5669\n" - ] - } - ], - "source": [ - "# Find recipes using the output embeddings model\n", - "top_k = 5\n", - "\n", - "# Find the most similar recipes to the given queries\n", - "query = \"cheeseburger\"\n", - "embeddings = model_custom.encode(recipes)\n", - "\n", - "results = find_most_similar_items(model_custom, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " \n", - "print()\n", - "\n", - "query = \"fattoush\"\n", - "results = find_most_similar_items(model_custom, embeddings, query, top_k)\n", - "print(f\"Most similar recipes to '{query}':\")\n", - "for idx, score in results:\n", - " print(f\"Title: `{recipes[idx]}`, Similarity Score: {score:.4f}\")\n", - " " - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "As can be seen, we now get good results for both queries with our custom vocab model since the domain-specific terms are included in the vocab." - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": "venv", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.10.11" - } - }, - "nbformat": 4, - "nbformat_minor": 2 -} diff --git a/tutorials/semantic_chunking.ipynb b/tutorials/semantic_chunking.ipynb deleted file mode 100644 index c18d182e..00000000 --- a/tutorials/semantic_chunking.ipynb +++ /dev/null @@ -1,268 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Semantic Chunking with Chonkie and Model2Vec**\n", - "\n", - "Semantic chunking is a task of identifying the semantic boundaries of a piece of text. In this tutorial, we will use the [Chonkie](https://github.com/bhavnicksm/chonkie) library to perform semantic chunking on the book War and Peace. Chonkie is a library that provides a lightweight and fast solution to semantic chunking using pre-trained models. It supports our [potion models](https://huggingface.co/collections/minishlab/potion-6721e0abd4ea41881417f062) out of the box, which we will be using in this tutorial.\n", - "\n", - "After chunking our text, we will be using [Vicinity](https://github.com/MinishLab/vicinity), a lightweight nearest neighbors library, to create an index of our chunks and query them." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "# Install the necessary libraries\n", - "!pip install -q datasets model2vec numpy tqdm vicinity \"chonkie[semantic]\"" - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "metadata": {}, - "outputs": [], - "source": [ - "# Import the necessary libraries\n", - "import random \n", - "import re\n", - "import requests\n", - "from time import perf_counter\n", - "from chonkie import SDPMChunker\n", - "from model2vec import StaticModel\n", - "from vicinity import Vicinity\n", - "\n", - "random.seed(0)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Loading and pre-processing**\n", - "\n", - "First, we will download War and Peace and apply some basic pre-processing." - ] - }, - { - "cell_type": "code", - "execution_count": 2, - "metadata": {}, - "outputs": [], - "source": [ - "# URL for War and Peace on Project Gutenberg\n", - "url = \"https://www.gutenberg.org/files/2600/2600-0.txt\"\n", - "\n", - "# Download the book\n", - "response = requests.get(url)\n", - "book_text = response.text\n", - "\n", - "def preprocess_text(text: str, min_length: int = 5):\n", - " \"\"\"Basic text preprocessing function.\"\"\"\n", - " text = text.replace(\"\\n\", \" \")\n", - " text = text.replace(\"\\r\", \" \")\n", - " sentences = re.findall(r'[^.!?]*[.!?]', text)\n", - " # Filter out sentences shorter than the specified minimum length\n", - " filtered_sentences = [sentence.strip() for sentence in sentences if len(sentence.split()) >= min_length]\n", - " # Recombine the filtered sentences\n", - " return ' '.join(filtered_sentences)\n", - "\n", - "# Preprocess the text\n", - "book_text = preprocess_text(book_text)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "**Chunking with Chonkie**\n", - "\n", - "Next, we will use Chonkie to chunk our text into semantic chunks." - ] - }, - { - "cell_type": "code", - "execution_count": 22, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Number of chunks: 4436\n", - "Time taken: 1.6311538339941762\n" - ] - } - ], - "source": [ - "# Initialize a SemanticChunker from Chonkie with the potion-base-8M model\n", - "chunker = SDPMChunker(\n", - " embedding_model=\"minishlab/potion-base-32M\",\n", - " chunk_size = 512, \n", - " skip_window=5, \n", - " min_sentences=3\n", - ")\n", - "\n", - "# Chunk the text\n", - "time = perf_counter()\n", - "chunks = chunker.chunk(book_text)\n", - "print(f\"Number of chunks: {len(chunks)}\")\n", - "print(f\"Time taken: {perf_counter() - time}\")" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "And that's it, we chunked the entirety of War and Peace in ~2 seconds. Not bad! Let's look at some example chunks." - ] - }, - { - "cell_type": "code", - "execution_count": 23, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - " Wait and we shall see! As if fighting were fun. They are like children from whom one can’t get any sensible account of what has happened because they all want to show how well they can fight. But that’s not what is needed now. “And what ingenious maneuvers they all propose to me! \n", - "\n", - " The first thing he saw on riding up to the space where TĂșshin’s guns were stationed was an unharnessed horse with a broken leg, that lay screaming piteously beside the harnessed horses. Blood was gushing from its leg as from a spring. Among the limbers lay several dead men. \n", - "\n", - " Out of an army of a hundred thousand we must expect at least twenty thousand wounded, and we haven’t stretchers, or bunks, or dressers, or doctors enough for six thousand. We have ten thousand carts, but we need other things as well—we must manage as best we can! ” The strange thought that of the thousands of men, young and old, who had stared with merry surprise at his hat (perhaps the very men he had noticed), twenty thousand were inevitably doomed to wounds and death amazed Pierre. “They may die tomorrow; why are they thinking of anything but death? ” And by some latent sequence of thought the descent of the MozhĂĄysk hill, the carts with the wounded, the ringing bells, the slanting rays of the sun, and the songs of the cavalrymen vividly recurred to his mind. “The cavalry ride to battle and meet the wounded and do not for a moment think of what awaits them, but pass by, winking at the wounded. Yet from among these men twenty thousand are doomed to die, and they wonder at my hat! ” thought Pierre, continuing his way to TatĂĄrinova. \n", - "\n" - ] - } - ], - "source": [ - "# Print a few example chunks\n", - "for _ in range(3):\n", - " chunk = random.choice(chunks)\n", - " print(chunk.text, \"\\n\")" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Those look good. Next, let's create a vector search index with Vicinity and Model2Vec.\n", - "\n", - "**Creating a vector search index**" - ] - }, - { - "cell_type": "code", - "execution_count": 24, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Time taken: 2.269912125004339\n" - ] - } - ], - "source": [ - "# Initialize an embedding model and encode the chunk texts\n", - "time = perf_counter()\n", - "model = StaticModel.from_pretrained(\"minishlab/potion-base-32M\")\n", - "chunk_texts = [chunk.text for chunk in chunks]\n", - "chunk_embeddings = model.encode(chunk_texts)\n", - "\n", - "# Create a Vicinity instance\n", - "vicinity = Vicinity.from_vectors_and_items(vectors=chunk_embeddings, items=chunk_texts)\n", - "print(f\"Time taken: {perf_counter() - time}\")" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Done! We embedded all our chunks and created an in index in ~1.5 seconds. Now that we have our index, let's query it with some queries.\n", - "\n", - "**Querying the index**" - ] - }, - { - "cell_type": "code", - "execution_count": 25, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Query: Emperor Napoleon\n", - "--------------------------------------------------\n", - " In 1808 the Emperor Alexander went to Erfurt for a fresh interview with the Emperor Napoleon, and in the upper circles of Petersburg there was much talk of the grandeur of this important meeting. CHAPTER XXII In 1809 the intimacy between “the world’s two arbiters,” as Napoleon and Alexander were called, was such that when Napoleon declared war on Austria a Russian corps crossed the frontier to co-operate with our old enemy Bonaparte against our old ally the Emperor of Austria, and in court circles the possibility of marriage between Napoleon and one of Alexander’s sisters was spoken of. But besides considerations of foreign policy, the attention of Russian society was at that time keenly directed on the internal changes that were being undertaken in all the departments of government. Life meanwhile—real life, with its essential interests of health and sickness, toil and rest, and its intellectual interests in thought, science, poetry, music, love, friendship, hatred, and passions—went on as usual, independently of and apart from political friendship or enmity with Napoleon Bonaparte and from all the schemes of reconstruction. BOOK SIX: 1808 - 10 CHAPTER I Prince Andrew had spent two years continuously in the country. All the plans Pierre had attempted on his estates—and constantly changing from one thing to another had never accomplished—were carried out by Prince Andrew without display and without perceptible difficulty. \n", - "\n", - " CHAPTER XXVI On August 25, the eve of the battle of BorodinĂł, M. de Beausset, prefect of the French Emperor’s palace, arrived at Napoleon’s quarters at ValĂșevo with Colonel Fabvier, the former from Paris and the latter from Madrid. Donning his court uniform, M. de Beausset ordered a box he had brought for the Emperor to be carried before him and entered the first compartment of Napoleon’s tent, where he began opening the box while conversing with Napoleon’s aides-de-camp who surrounded him. Fabvier, not entering the tent, remained at the entrance talking to some generals of his acquaintance. The Emperor Napoleon had not yet left his bedroom and was finishing his toilet. \n", - "\n", - " In Russia there was an Emperor, Alexander, who decided to restore order in Europe and therefore fought against Napoleon. In 1807 he suddenly made friends with him, but in 1811 they again quarreled and again began killing many people. Napoleon led six hundred thousand men into Russia and captured Moscow; then he suddenly ran away from Moscow, and the Emperor Alexander, helped by the advice of Stein and others, united Europe to arm against the disturber of its peace. All Napoleon’s allies suddenly became his enemies and their forces advanced against the fresh forces he raised. The Allies defeated Napoleon, entered Paris, forced Napoleon to abdicate, and sent him to the island of Elba, not depriving him of the title of Emperor and showing him every respect, though five years before and one year later they all regarded him as an outlaw and a brigand. Then Louis XVIII, who till then had been the laughingstock both of the French and the Allies, began to reign. And Napoleon, shedding tears before his Old Guards, renounced the throne and went into exile. \n", - "\n", - "Query: The battle of Austerlitz\n", - "--------------------------------------------------\n", - " Behave as you did at Austerlitz, Friedland, VĂ­tebsk, and SmolĂ©nsk. Let our remotest posterity recall your achievements this day with pride. Let it be said of each of you: “He was in the great battle before Moscow! \n", - "\n", - " By a strange coincidence, this task, which turned out to be a most difficult and important one, was entrusted to DokhtĂșrov—that same modest little DokhtĂșrov whom no one had described to us as drawing up plans of battles, dashing about in front of regiments, showering crosses on batteries, and so on, and who was thought to be and was spoken of as undecided and undiscerning—but whom we find commanding wherever the position was most difficult all through the Russo-French wars from Austerlitz to the year 1813. At Austerlitz he remained last at the Augezd dam, rallying the regiments, saving what was possible when all were flying and perishing and not a single general was left in the rear guard. Ill with fever he went to SmolĂ©nsk with twenty thousand men to defend the town against Napoleon’s whole army. \n", - "\n", - " “Nothing is truer or sadder. These gentlemen ride onto the bridge alone and wave white handkerchiefs; they assure the officer on duty that they, the marshals, are on their way to negotiate with Prince Auersperg. He lets them enter the tĂȘte-de-pont. * They spin him a thousand gasconades, saying that the war is over, that the Emperor Francis is arranging a meeting with Bonaparte, that they desire to see Prince Auersperg, and so on. The officer sends for Auersperg; these gentlemen embrace the officers, crack jokes, sit on the cannon, and meanwhile a French battalion gets to the bridge unobserved, flings the bags of incendiary material into the water, and approaches the tĂȘte-de-pont. At length appears the lieutenant general, our dear Prince Auersperg von Mautern himself. Flower of the Austrian army, hero of the Turkish wars! Hostilities are ended, we can shake one another’s hand. The Emperor Napoleon burns with impatience to make Prince Auersperg’s acquaintance. \n", - "\n", - "Query: Paris\n", - "--------------------------------------------------\n", - " Paris is Talma, la DuchĂ©nois, Potier, the Sorbonne, the boulevards,” and noticing that his conclusion was weaker than what had gone before, he added quickly: “There is only one Paris in the world. You have been to Paris and have remained Russian. Well, I don’t esteem you the less for it. \n", - "\n", - " Look at our youths, look at our ladies! The French are our Gods: Paris is our Kingdom of Heaven. ” He began speaking louder, evidently to be heard by everyone. “French dresses, French ideas, French feelings! \n", - "\n", - " “Oh yes, one sees that plainly. A man who doesn’t know Paris is a savage. You can tell a Parisian two leagues off. \n", - "\n" - ] - } - ], - "source": [ - "queries = [\"Emperor Napoleon\", \"The battle of Austerlitz\", \"Paris\"]\n", - "for query in queries:\n", - " print(f\"Query: {query}\\n{'-' * 50}\")\n", - " query_embedding = model.encode(query)\n", - " results = vicinity.query(query_embedding, k=3)[0]\n", - "\n", - " for result in results:\n", - " print(result[0], \"\\n\")" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "These indeed look like relevant chunks, nice! That's it for this tutorial. We were able to chunk, index, and query War and Peace in about 3.5 seconds using Chonkie, Vicinity, and Model2Vec. Lightweight and fast, just how we like it." - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": "Python 3 (ipykernel)", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.10.15" - } - }, - "nbformat": 4, - "nbformat_minor": 4 -} diff --git a/tutorials/train_classifier.ipynb b/tutorials/train_classifier.ipynb deleted file mode 100644 index 988007d6..00000000 --- a/tutorials/train_classifier.ipynb +++ /dev/null @@ -1,806 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "# Training a classifier using model2vec\n", - "\n", - "Model2Vec supports built-in classifier training with an easy, scikit-learn-based syntax. Just give the model your data in `.fit`, and you'll have a trained model!\n", - "\n", - "How it works:\n", - "* We load a base `StaticModel` using as a torch module. By default we use [potion-base-8m](https://huggingface.co/minishlab/potion-base-8M).\n", - "* We add a one-layer MLP with 512 hidden units and `ReLU` activation as a head.\n", - "* We train the model using cross-entropy, using [`pytorch-lightning`](https://lightning.ai/docs/pytorch/stable/) as a training framework.\n", - "\n", - "After training, you can export the model using regular torch tools, such as `torch.save` and `torch.load`, or you can export the model to a `scikit-learn` pipeline. The latter option leads to a really small footprint during inference, as there is no longer a need to use `torch`." - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "metadata": { - "vscode": { - "languageId": "plaintext" - } - }, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "\u001b[2mUsing Python 3.11.4 environment at: /Users/stephantulkens/Documents/GitHub/model2vec/.venv\u001b[0m\n", - "\u001b[2mAudited \u001b[1m1 package\u001b[0m \u001b[2min 4ms\u001b[0m\u001b[0m\n", - "\u001b[2mUsing Python 3.11.4 environment at: /Users/stephantulkens/Documents/GitHub/model2vec/.venv\u001b[0m\n", - "\u001b[2mAudited \u001b[1m1 package\u001b[0m \u001b[2min 8ms\u001b[0m\u001b[0m\n", - "\u001b[2mUsing Python 3.11.4 environment at: /Users/stephantulkens/Documents/GitHub/model2vec/.venv\u001b[0m\n", - "\u001b[2mAudited \u001b[1m1 package\u001b[0m \u001b[2min 3ms\u001b[0m\u001b[0m\n" - ] - } - ], - "source": [ - "# Install the necessary libraries\n", - "!uv pip install \"model2vec[train,inference]\"\n", - "!uv pip install \"datasets\"\n", - "!uv pip install \"scikit-learn\"\n", - "\n", - "# Import the necessary libraries\n", - "from model2vec.train import StaticModelForClassification\n", - "from model2vec.inference import StaticModelPipeline" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "To demonstrate how to train a model, we'll be using the `20_newsgroups` dataset, which contains posts from 1 of 20 newsgroups." - ] - }, - { - "cell_type": "code", - "execution_count": 2, - "metadata": {}, - "outputs": [ - { - "name": "stderr", - "output_type": "stream", - "text": [ - "Repo card metadata block was not found. Setting CardData to empty.\n" - ] - }, - { - "name": "stdout", - "output_type": "stream", - "text": [ - "DatasetDict({\n", - " train: Dataset({\n", - " features: ['text', 'label', 'label_text'],\n", - " num_rows: 11314\n", - " })\n", - " test: Dataset({\n", - " features: ['text', 'label', 'label_text'],\n", - " num_rows: 7532\n", - " })\n", - "})\n" - ] - } - ], - "source": [ - "from datasets import load_dataset\n", - "\n", - "dataset = load_dataset(\"setfit/20_newsgroups\")\n", - "print(dataset)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Let's take a look at the first five training samples:" - ] - }, - { - "cell_type": "code", - "execution_count": 3, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "TEXT: I was wondering if anyone out there could enlighten me on this car I saw\n", - "the other day. It was a 2-door sports car, looked to be from the late 60s/\n", - "early 70s. It was called a Bricklin. The doors were really small. In addition,\n", - "the front bumper was separate from the rest of the body. This is \n", - "all I know. If anyone can tellme a model name, engine specs, years\n", - "of production, where this car is made, history, or whatever info you\n", - "have on this funky looking car, please e-mail. LABEL: rec.autos\n", - "TEXT: A fair number of brave souls who upgraded their SI clock oscillator have\n", - "shared their experiences for this poll. Please send a brief message detailing\n", - "your experiences with the procedure. Top speed attained, CPU rated speed,\n", - "add on cards and adapters, heat sinks, hour of usage per day, floppy disk\n", - "functionality with 800 and 1.4 m floppies are especially requested.\n", - "\n", - "I will be summarizing in the next two days, so please add to the network\n", - "knowledge base if you have done the clock upgrade and haven't answered this\n", - "poll. Thanks. LABEL: comp.sys.mac.hardware\n", - "TEXT: well folks, my mac plus finally gave up the ghost this weekend after\n", - "starting life as a 512k way back in 1985. sooo, i'm in the market for a\n", - "new machine a bit sooner than i intended to be...\n", - "\n", - "i'm looking into picking up a powerbook 160 or maybe 180 and have a bunch\n", - "of questions that (hopefully) somebody can answer:\n", - "\n", - "* does anybody know any dirt on when the next round of powerbook\n", - "introductions are expected? i'd heard the 185c was supposed to make an\n", - "appearence \"this summer\" but haven't heard anymore on it - and since i\n", - "don't have access to macleak, i was wondering if anybody out there had\n", - "more info...\n", - "\n", - "* has anybody heard rumors about price drops to the powerbook line like the\n", - "ones the duo's just went through recently?\n", - "\n", - "* what's the impression of the display on the 180? i could probably swing\n", - "a 180 if i got the 80Mb disk rather than the 120, but i don't really have\n", - "a feel for how much \"better\" the display is (yea, it looks great in the\n", - "store, but is that all \"wow\" or is it really that good?). could i solicit\n", - "some opinions of people who use the 160 and 180 day-to-day on if its worth\n", - "taking the disk size and money hit to get the active display? (i realize\n", - "this is a real subjective question, but i've only played around with the\n", - "machines in a computer store breifly and figured the opinions of somebody\n", - "who actually uses the machine daily might prove helpful).\n", - "\n", - "* how well does hellcats perform? ;)\n", - "\n", - "thanks a bunch in advance for any info - if you could email, i'll post a\n", - "summary (news reading time is at a premium with finals just around the\n", - "corner... :( )\n", - "--\n", - "Tom Willis \\ twillis@ecn.purdue.edu \\ Purdue Electrical Engineering LABEL: comp.sys.mac.hardware\n", - "TEXT: \n", - "Do you have Weitek's address/phone number? I'd like to get some information\n", - "about this chip.\n", - " LABEL: comp.graphics\n", - "TEXT: From article , by tombaker@world.std.com (Tom A Baker):\n", - "\n", - "\n", - "My understanding is that the 'expected errors' are basically\n", - "known bugs in the warning system software - things are checked\n", - "that don't have the right values in yet because they aren't\n", - "set till after launch, and suchlike. Rather than fix the code\n", - "and possibly introduce new bugs, they just tell the crew\n", - "'ok, if you see a warning no. 213 before liftoff, ignore it'. LABEL: sci.space\n" - ] - } - ], - "source": [ - "# First 5 training samples:\n", - "for record in dataset[\"train\"].to_list()[:5]:\n", - " print(f\"TEXT: {record['text']} LABEL: {record['label_text']}\")" - ] - }, - { - "cell_type": "code", - "execution_count": 4, - "metadata": {}, - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "StaticModelForClassification(\n", - " (embeddings): Embedding(29528, 256, padding_idx=0)\n", - " (head): Sequential(\n", - " (0): Linear(in_features=256, out_features=512, bias=True)\n", - " (1): ReLU()\n", - " (2): Linear(in_features=512, out_features=2, bias=True)\n", - " )\n", - ")\n" - ] - } - ], - "source": [ - "# Define the staticmodel\n", - "model = StaticModelForClassification.from_pretrained()\n", - "# Optional arguments:\n", - "# model_name: the name of the base model (defaults to potion-base-8m)\n", - "# n_layers: the number of layers in the MLP (defaults to 1)\n", - "# hidden_dim: the number of hidden units (defaults to 512)\n", - "print(model)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Now let's train the model on a subset of examples. We pick the first 1000 examples to train on." - ] - }, - { - "cell_type": "code", - "execution_count": 5, - "metadata": {}, - "outputs": [ - { - "name": "stderr", - "output_type": "stream", - "text": [ - "Seed set to 42\n", - "GPU available: True (mps), used: True\n", - "TPU available: False, using: 0 TPU cores\n", - "HPU available: False, using: 0 HPUs\n", - "/Users/stephantulkens/Documents/GitHub/model2vec/.venv/lib/python3.11/site-packages/lightning/pytorch/trainer/connectors/logger_connector/logger_connector.py:76: Starting from v1.9.0, `tensorboardX` has been removed as a dependency of the `lightning.pytorch` package, due to potential conflicts with other packages in the ML ecosystem. For this reason, `logger=True` will use `CSVLogger` as the default logger, unless the `tensorboard` or `tensorboardX` packages are found. Please `pip install lightning[extra]` or one of them to enable TensorBoard support by default\n", - "/Users/stephantulkens/Documents/GitHub/model2vec/.venv/lib/python3.11/site-packages/torch/optim/lr_scheduler.py:60: UserWarning: The verbose parameter is deprecated. Please use get_last_lr() to access the learning rate.\n", - " warnings.warn(\n", - "\n", - " | Name | Type | Params | Mode \n", - "---------------------------------------------------------------\n", - "0 | model | StaticModelForClassification | 7.7 M | train\n", - "---------------------------------------------------------------\n", - "7.7 M Trainable params\n", - "0 Non-trainable params\n", - "7.7 M Total params\n", - "30.922 Total estimated model params size (MB)\n", - "6 Modules in train mode\n", - "0 Modules in eval mode\n" - ] - }, - { - "data": { - "application/vnd.jupyter.widget-view+json": { - "model_id": "2351ba8c0b53458fb680e8d29e0f0a6c", - "version_major": 2, - "version_minor": 0 - }, - "text/plain": [ - "Sanity Checking: | | 0/? [00:00