Skip to main content
PineconeBuying intent

A company is considering migrating its database infrastructure to Pinecone to overcome the vector dimension limits of its current pgvector setup and support Google's Gemini embeddings.

What happened

Post: "Migrate the embeddings model or the Database infrastructure"

Source

RedditSep 25, 2026By u/RecentTheory6635

r/Rag

Migrate the embeddings model or the Database infrastructure

upvotes
5
comments
9

Post

Highlighted: the lines this signal was extracted from

Hey! Im usually active on these feeds but never comment but this problems made me stress out. I've built a rag across 2000 documents for my company right now and i was originally using the Text-embeddings-3-model from OpenAI and realized it wasn't able to gather nuanced contexts like images so i decided to migrate to the Gemini multi modal. Our current DB runs on supabase and uses pg vector as our vector db. Currently we use HNSW search + BM 25 in our search algorithm and hit a constraint during migration as Geminis vectors are bigger than the 2000 limit we get using postgreSQL. We can either use truncated vectors and accept some loss of information or migrate to a cohereV4 multimodal tech that fits in the vector constraints allowed. My coworker wants to migrate our entire DB to something like pinecone but something tells me migrating Databases for something like this isn't worth doing. (We aren't in production for other users yet and have a small blast radius). I would love to know your suggestions!

reddit.com/r/Rag/comments/1wq1tfx/migrate_the_embeddings_model_or_the...Read the full source

Comments on the post

5 of 9 comments
  • “With only 2k docs, I'd first test the truncated Gemini vectors and see how much retrieval quality drops. You already have BM25 + HNSW, so it may be fine. If the dimension limit is still an issue, I'd use a separate vector search engine instead of moving the whole DB. Keep Supabase for your main data. I'm building one myself. It's been tested on 10M Cohere Wikipedia vectors (1024D, 41GB raw dat”

    u/24iqSuperGenius1 points · Sep 25, 2026View

  • “You can use https://index-server.searchblox.com/blog/semantic-search-dialects.html and https://inference-server.searchblox.com/blog/multimodal-embeddings-reranker-local.html for free and run locally.”

    u/searchblox_searchai1 points · Sep 25, 2026View

  • “Before changing databases, I’d separate missing evidence from poor retrieval. For a few failed image-dependent questions, check whether the required visual content reaches the embedding step at all. Then compare embedding options against the same questions and expected evidence. A database migration won’t help if the relevant information never entered the index.”

    u/OkFlan5041 points · Sep 25, 2026View

  • “I’d be hesitant to move the whole database before proving the larger vectors are buying you enough quality to justify it. With 2k docs and no external users yet, this feels like a good point to build a small eval set from your harder queries and test the truncated Gemini vectors against the other embedding option first. A DB migration is a lot easier to justify when you can point to failures the”

    u/Clean_Research_25831 points · Sep 25, 2026View

  • “I built www.embedding-adapters.com for this reason”

    u/Interesting-Town-4331 points · Sep 25, 2026View

Extracted by Autobound

From the Signal API record
Signal
Buying intent

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/Rag
Stage
Evaluating

Companies

  • SupabaseAlso named

The full record

From the Signal API record

Numbers

Mentions
2

Details

Timing
In progress
Category
Vector Database
Virality
Medium
Post kind
Text
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • database
  • vector db
  • migration
  • vendor evaluation
  • rag

Flair

  • Discussion

Extraction

Sentiment
Neutral
Detected
Sep 25, 2026
signal_type
reddit-company
signal_subtype
buyingIntent

Use this data

Get every Reddit signal for Pinecone and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Pinecone this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/a64049cf-27f1-5bfb-a762-b8b05ae0e9dd returns this record as JSON. POST /v1/companies/enrich returns every signal for pinecone.io.

{
  "signal_id": "a64049cf-27f1-5bfb-a762-b8b05ae0e9dd",
  "signal_type": "reddit-company",
  "signal_subtype": "buyingIntent",
  "detected_at": "2026-09-25T17:04:23+00:00",
  "company": {
    "name": "Pinecone",
    "domain": "pinecone.io"
  },
  "data": {
    "nsfw": false,
    "stage": "evaluating",
    "awards": 0,
    "timing": "in_progress",
    "topics": [
      "database",
      "vector db",
      "rag",
      "migration",
      "vendor evaluation"
    ],
    "post_id": "1wq1tfx",
    "summary": "A company is considering migrating its database infrastructure to Pinecone to overcome the vector dimension limits of its current pgvector setup and support Google's Gemini embeddings.",
    "category": "Vector Database",
    "comments": [
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc1fsl0/",
        "depth": 0,
        "score": 1,
        "author": "24iqSuperGenius",
        "excerpt": "With only 2k docs, I'd first test the truncated Gemini vectors and see how much retrieval quality drops. You already have BM25 + HNSW, so it may be fine.\n\n If the dimension limit is still an issue, I'd use a separate vector search engine instead of moving the whole DB. Keep Supabase for your main data.\n\n I'm building one myself. It's been tested on 10M Cohere Wikipedia vectors (1024D, 41GB raw dat",
        "posted_at": "2026-09-25T20:05:04.000Z",
        "author_url": "https://www.reddit.com/user/24iqSuperGenius/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc1kyv4/",
        "depth": 0,
        "score": 1,
        "author": "searchblox_searchai",
        "excerpt": "You can use https://index-server.searchblox.com/blog/semantic-search-dialects.html and https://inference-server.searchblox.com/blog/multimodal-embeddings-reranker-local.html for free and run locally.",
        "posted_at": "2026-09-25T20:28:04.000Z",
        "author_url": "https://www.reddit.com/user/searchblox_searchai/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc1sdvg/",
        "depth": 0,
        "score": 1,
        "author": "OkFlan504",
        "excerpt": "Before changing databases, I’d separate missing evidence from poor retrieval. For a few failed image-dependent questions, check whether the required visual content reaches the embedding step at all. Then compare embedding options against the same questions and expected evidence. A database migration won’t help if the relevant information never entered the index.",
        "posted_at": "2026-09-25T21:01:21.000Z",
        "author_url": "https://www.reddit.com/user/OkFlan504/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc29xsr/",
        "depth": 0,
        "score": 1,
        "author": "Clean_Research_2583",
        "excerpt": "I’d be hesitant to move the whole database before proving the larger vectors are buying you enough quality to justify it. With 2k docs and no external users yet, this feels like a good point to build a small eval set from your harder queries and test the truncated Gemini vectors against the other embedding option first.\n\n A DB migration is a lot easier to justify when you can point to failures the",
        "posted_at": "2026-09-25T22:27:17.000Z",
        "author_url": "https://www.reddit.com/user/Clean_Research_2583/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc2p8eb/",
        "depth": 0,
        "score": 1,
        "author": "Interesting-Town-433",
        "excerpt": "I built www.embedding-adapters.com for this reason",
        "posted_at": "2026-09-25T23:45:00.000Z",
        "author_url": "https://www.reddit.com/user/Interesting-Town-433/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pcdvzth/",
        "depth": 0,
        "score": 1,
        "author": "Lumpy_Fisherman_5352",
        "excerpt": "Migrating to Pinecone for 2000 docs is overkill - stay on pgvector. Two real options: (1) pgvector 0.7+ has halfvec, which indexes up to 4000 dims, so a 3072-dim Gemini vector fits with HNSW at half the storage; or (2) both Gemini and OpenAI-3 are Matryoshka-trained, so setting a smaller output dimensionality (e.g. 1536) is a designed truncation, not the lossy kind - usually a tiny recall hit. I'd",
        "posted_at": "2026-09-27T15:42:46.000Z",
        "author_url": "https://www.reddit.com/user/Lumpy_Fisherman_5352/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/comment/pc0wpeo/",
        "depth": 0,
        "score": 1,
        "author": "onur90",
        "excerpt": "Why don't you try it with a graph database (Neo4j)? I installed Neo4j with the MCP server. The RAG can use MCP and Cypher (it's like SQL for graph databases).\nThe RAG can perform similarity searches to find the first node in the graph. From there, it searches node by node until it finds the information it needs.\nTo create the graph, I used a custom script with Gemini AI.",
        "posted_at": "2026-09-25T18:40:51.000Z",
        "author_url": "https://www.reddit.com/user/onur90/"
      }
    ],
    "evidence": [
      "[post] My coworker wants to migrate our entire DB to something like pinecone but something tells me migrating Databases for something like this isn't worth doing."
    ],
    "virality": "medium",
    "post_date": "2026-09-25T17:04:23.000Z",
    "post_kind": "text",
    "post_text": "Hey! Im usually active on these feeds but never comment but this problems made me stress out.\n\nI've built a rag across 2000 documents for my company right now and i was originally using the Text-embeddings-3-model from OpenAI and realized it wasn't able to gather nuanced contexts like images so i decided to migrate to the Gemini multi modal.\n\nOur current DB runs on supabase and uses pg vector as our vector db. Currently we use HNSW search + BM 25 in our search algorithm and hit a constraint during migration as Geminis vectors are bigger than the 2000 limit we get using postgreSQL.\n\nWe can either use truncated vectors and accept some loss of information or migrate to a cohereV4 multimodal tech that fits in the vector constraints allowed. My coworker wants to migrate our entire DB to something like pinecone but something tells me migrating Databases for something like this isn't worth doing. (We aren't in production for other users yet and have a small blast radius).\n\nI would love to know your suggestions!",
    "sentiment": "neutral",
    "subreddit": "Rag",
    "post_flair": [
      "Discussion"
    ],
    "post_title": "Migrate the embeddings model or the Database infrastructure",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/Rag/comments/1wq1tfx/migrate_the_embeddings_model_or_the_database/",
    "entity_role": "vendor",
    "post_author": "RecentTheory6635",
    "upvote_ratio": 0.8571428571428571,
    "mention_count": 2,
    "mention_surge": false,
    "subreddit_url": "https://www.reddit.com/r/Rag/",
    "total_upvotes": 5,
    "comments_total": 9,
    "total_comments": 9,
    "other_companies": [
      {
        "name": "Supabase",
        "role": "incumbent",
        "domain": "supabase.com"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/RecentTheory6635/",
    "signal_category": "intent",
    "comments_included": 7
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.