Is the fastest vector index always the right vector index? HNSW is difficult to beat when an application needs the lowest latency and highest query throughput. But production systems optimise for different constraints: 1. Some need to support very large datasets without keeping the full index in memory. 2. Some are limited by available memory or infrastructure cost. 3. Some change frequently and cannot depend on disruptive full-index rebuilds. HFresh takes a different approach. It uses a compact HNSW index to route each query towards the most relevant centroids, then searches compressed postings stored on disk. Underneath that architecture: - RQ8 makes centroid vectors approximately four times smaller before overhead. - RQ1 reduces stored posting-vector data by up to 32 times compared with 32-bit floats. - Background split, merge, and reassign operations keep the index balanced incrementally. The result is a disk-based index designed for lower heap usage, controlled query I/O, and large mutable datasets—without periodic full rebuilds. The choice is not HFresh or HNSW in every situation. Choose HNSW when peak query performance is the priority. Choose HFresh when memory efficiency matters more than peak throughput. Read how HFresh works: https://lnkd.in/dgPHmXXg
Weaviate
Technology, Information and Internet
Amsterdam, North Holland 59,439 followers
The AI database for a new generation of software.
About us
Weaviate is a cloud-native, real-time vector database that allows you to bring your machine-learning models to scale. There are extensions for specific use cases, such as semantic search, plugins to integrate Weaviate in any application of your choice, and a console to visualize your data.
- Website
-
https://weaviate.io
External link for Weaviate
- Industry
- Technology, Information and Internet
- Company size
- 51-200 employees
- Headquarters
- Amsterdam, North Holland
- Type
- Privately Held
- Founded
- 2019
Locations
-
Primary
Get directions
Amsterdam, North Holland, NL
Employees at Weaviate
Updates
-
What if you remember exactly what an asset looks like, but not what it was called? Creative teams rarely think in filenames. They remember a mood, a scene, a character, or a half-finished idea from an earlier project. The archive remembers something else: `final_FINAL_v7.svg` `BROLL_NEW2.svg` `scene_14_USE_THIS.svg` In the final episode of Building Foundry, we turn that gap into a working search experience. We explore: 1. How to scan a creative archive without moving or renaming the original files. 2. How descriptions, tags, extracted text, and metadata give each asset useful context. 3. How keyword, semantic, and hybrid search support different ways of remembering. The result is a searchable layer over the archive teams already have. A search for “white rabbit in a grassy landscape” can retrieve a reference image, a trailer clip, and related landscape assets, even when their filenames describe none of those things. Read the final episode: https://lnkd.in/diHehtcF
-
-
Weaviate reposted this
Your agent forgets what you told it 10 minutes ago. Or worse, it retrieves "memories" that have nothing to do with the current task. Most agentic memory systems just store every message, search that pile with the latest user turn, and dump the results into context. We've spent the last year building a memory service at Weaviate, and that approach breaks in three specific ways. 𝗖𝗼𝘀𝘁 𝗮𝗻𝗱 𝗰𝗼𝗻𝘁𝗲𝘅𝘁 𝗴𝗿𝗼𝘄𝘁𝗵. You're paying for it on every single message. Your entire conversation history from every chat compounds over time, eats tokens, slows the model down, and makes answers worse. 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗶𝘀 𝗿𝗲𝗮𝗰𝘁𝗶𝘃𝗲. What you get back is only as good as the text you searched with - which is usually the last user message, and often vague. 𝗖𝗼𝗻𝘁𝗿𝗮𝗱𝗶𝗰𝘁𝗶𝗼𝗻 𝗮𝗻𝗱 𝘀𝘁𝗮𝗹𝗲𝗻𝗲𝘀𝘀. This is the one we spent the longest engineering around. People change their minds. Facts go stale. Extract information from every message independently and you end up with 50 memories about the same thing, all slightly different, all in the context window. Dumping duplicates, contradictions, and outdated facts into context means asking the model to sort it out at query time. That's the worst possible moment to give it extra work, and every conflicting piece makes the response a little worse. So 𝗘𝗻𝗴𝗿𝗮𝗺 treats memory as engineered infrastructure instead of storage bolted onto the app. Memory gets 𝗮𝗰𝘁𝗶𝘃𝗲𝗹𝘆 𝗺𝗮𝗶𝗻𝘁𝗮𝗶𝗻𝗲𝗱, not accumulated. New video walking through why these failures happen, plus a full setup of a basic chat app with Engram 👇 https://lnkd.in/dBZwvuD8 Learn more about Engram: https://lnkd.in/d_gSsEP3 Demo: https://lnkd.in/dUDut_aD
-
Weaviate reposted this
Who knows... One day you might: https://lnkd.in/eHddJts5
-
-
Weaviate reposted this
💡 Did you know that Weaviate supports MCP (Model Context Protocol) as a native endpoint? 🧠 Unlike external MCP-wrappers, Weaviate supports MCP natively right in the core engine. That means zero middleware to deploy, lower query latency, and unified API key authentication straight to your DB 🔒 It stays secure using standard bearer token (API key) authentication, covering both your core Weaviate instance and MCP access out of the box ☁️ Directly accessible in the Weaviate Cloud (image below)! 👉 Test it here: https://lnkd.in/eySxPQFu
-
-
Filters exclude. Boost reorders. New in Weaviate 1.38, the 𝗕𝗼𝗼𝘀𝘁 𝗔𝗣𝗜 adds query-time rescoring when a hard filter is too strict. Take a search for 𝘆𝗲𝗹𝗹𝗼𝘄 𝗮𝗿𝗺𝗰𝗵𝗮𝗶𝗿: Filtering on 𝗶𝗻_𝘀𝘁𝗼𝗰𝗸=𝘁𝗿𝘂𝗲 removes every out-of-stock result. Boosting the same condition ranks available products higher while keeping other relevant matches in the response. Boost runs after the primary search and supports: • filter conditions • numeric property values • time decay for recency • numeric decay around a target value You can blend up to 20 conditions, adjust how strongly they influence the original relevance score, or use negative condition weights to demote matches without excluding them. Boost only reorders candidates retrieved by the primary search. The 𝗱𝗲𝗽𝘁𝗵 parameter controls the size of that candidate pool. Learn more: https://lnkd.in/euDXesc9
-
-
Skip the evaluation and get a faster, cheaper answer. Or get a precise set of sources that match your answer. Now you have full control between latency and verifiability. 𝗔𝘀𝗸 𝗠𝗼𝗱𝗲 turns a natural-language question into searches or aggregations across your Weaviate Cloud collections, then generates an answer grounded in the retrieved data. Previously, the result evaluation step always ran. Now you can customise the behavior of your ask mode run, using 𝗿𝗲𝘀𝘂𝗹𝘁_𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 in Python or 𝗿𝗲𝘀𝘂𝗹𝘁𝗘𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 in JavaScript/TypeScript: - Set it to "𝘯𝘰𝘯𝘦" to retrieve data and generate the final answer. This is the default, faster and cheaper workflow. - Set it to "𝘭𝘭𝘮" to evaluate the answer against the retrieved context using an additional model call. This retains the sources relevant to or used in the answer, and report possible incompleteness or missing information. The workflow becomes: 1. Retrieve data 2. Generate an answer 3. Optionally evaluate the answer and refine its sources Since the evaluation requires an additional LLM call, it adds latency and cost. But it also enables extra fields: 𝘪𝘴_𝘱𝘢𝘳𝘵𝘪𝘢𝘭_𝘢𝘯𝘴𝘸𝘦𝘳 and 𝘮𝘪𝘴𝘴𝘪𝘯𝘨_𝘪𝘯𝘧𝘰𝘳𝘮𝘢𝘵𝘪𝘰𝘯, giving you useful signals about whether the answer is complete and what may still be missing. You can now choose the right tradeoff for each application: faster responses when source refinement isn't required, or additional evaluation when attribution and completeness signals matter. Read the Ask Mode documentation: https://lnkd.in/e8gaUS8K
-
-
We stopped asking PDFs to become text before we searched them. With late-interaction multi-vector retrieval, you can embed each PDF page as an image and search the charts, tables, and layouts that text extraction misses. We tested it on 92 pages of NVIDIA investor decks. A query about automotive revenue retrieved the exact five-quarter bar chart, even though the words "change" and "over time" never appeared on the page. No OCR. No chunking. Just drag your PDFs into Weaviate Cloud and start asking questions. See how it works: https://lnkd.in/ebepfauR
-
-
Weaviate 1.39 is out 🚀 Two search features graduate to GA, plus a preview, one experimental, and a much quieter disk: 🎯 𝗕𝗼𝗼𝘀𝘁 𝗔𝗣𝗜 (𝗚𝗔) - nudge results up or down at query time. Nothing gets dropped, things just move in the results as you want them to. 👉 𝗠𝗠𝗥 𝗱𝗶𝘃𝗲𝗿𝘀𝗶𝘁𝘆 𝘀𝗲𝗹𝗲𝗰𝘁𝗶𝗼𝗻 (𝗚𝗔) - stop page one from being the same answer nine times. Works on hybrid and vector search. 🗜️ 𝟰-𝗯𝗶𝘁 𝗥𝗤 (𝗽𝗿𝗲𝘃𝗶𝗲𝘄) - If 8-bit compression was too much for you and 1-bit wasn’t enough, you will love this feature. 🌐 𝗦𝗲𝗮𝗿𝗰𝗵 𝗥𝗘𝗦𝗧 𝗔𝗣𝗜 (𝗲𝘅𝗽𝗲𝗿𝗶𝗺𝗲𝗻𝘁𝗮𝗹) - plain JSON over HTTP. No client, no gRPC, no GraphQL. Handy from a shell script or an edge function. 💾 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗶𝗰 𝗛𝗡𝗦𝗪 𝘀𝗻𝗮𝗽𝘀𝗵𝗼𝘁𝘀 - commit logs get cleaned up on their own. Less disk, faster startup, five settings you no longer need to think about. Also worth knowing: gRPC-Web has been on by default since 1.38.3, so you can reach the gRPC API straight from a browser. Open source on GitHub, or spin up a free cluster on Weaviate Cloud! Learn more: https://lnkd.in/eKy-24qe
-
-
RAG has a text-shaped bottleneck. Transcripts lose tone. OCR mangles layouts. Captions miss what happens on screen. Native multimodal RAG skips that text conversion step. With Google DeepMind's Gemini Embedding 2 and Weaviate, text, images, audio, and video can be embedded in one shared vector space. One query can retrieve a PDF page, an audio chunk, or the right moment in a video. Use it when the original media carries meaning that text cannot preserve. If your corpus is pure text, sticking with text embeddings is just fine. They are cheaper, faster, and usually enough. Explore the carousel and read the full guide with runnable code examples: https://lnkd.in/dR3n5i3g