Skip to content

EmbeddingGemma 2 ships under Apache 2.0: with embeddings, open weights are insurance, not a deployment plan

Published on 7 October 2026

Archivador lleno de fichas idénticas, una persona de espaldas iluminándolo con una linterna y una llave de repuesto colgada en primer plano.

Google has released EmbeddingGemma 2 under the Apache 2.0 licence, and the most useful comment about the launch isn't about benchmarks: it's about what happens the day the model disappears.

What happened

An embedding model turns a piece of text into a list of numbers — a vector — that lets you compare texts by meaning rather than by exact wording. The news here is that this version ships with an Apache 2.0 licence: downloadable weights, commercial use allowed.

Simon Willison commented on Hacker News and his argument is the one that matters. For embeddings specifically, betting on a closed model that only exists inside a vendor's API doesn't add up. These models aren't used once. They're used to compute thousands or millions of vectors that you then store for later comparison. When the vendor retires that model — and they will, because they'll have a better one — you don't lose a single API call: you lose the whole index, and you pay to rebuild it.

There's a precedent he cites himself: in April 2024, OpenAI said it would cover the cost of re-embedding content with the new models. A nice gesture, and exactly the sort of thing you can't assume from the next provider along.

And one nuance almost everyone drops when summarising him:

It's not that I want to host the model myself.

He'd rather pay someone to serve it, knowing that if they shut it down he can run the open weights himself or find someone else who will. Open weights as a policy, not as an architecture.

Why it matters if you build things

The difference between a chat model and an embedding model isn't technical, it's accounting. A chat response is ephemeral: you use it and bin it, and switching vendor means changing a string in your config. An embedding is persisted state. It lives in your database, and switching models means rebuilding that state in full.

Four things I'd do today, whichever model you use:

  1. Store the source text and the chunking, not just the vector. The expensive mistake isn't losing the model: it's having a million vectors and no copy of the text fragments that produced them, nor the code that split them that way. If you can't reproduce the chunking, you can't migrate even if the new model is free.
  2. Version the vector the way you version your schema. Model name, version, dimensions, whether you normalise, and which prompt template you used. Plenty of embedding models use different prefixes for queries and documents; mix them up and quality drops with nobody knowing why.
  3. Never mix vectors from two models in one index. The distance between an old-model vector and a new-model one means nothing. Migrate with a parallel index and dual writes, and cut over when the new one wins.
  4. Keep a set of real queries with the results you expect. Without that you can't tell whether you migrated well or quietly made search worse, which is what usually happens.

The cost of recomputing isn't the GPU bill. It's the reindexing window, the relevance regression, and everything hanging underneath: caches, summaries, whatever you built on top of those vectors.

What doesn't change

Apache 2.0 doesn't save you from migrating. When the next version lands and it's better, you'll want the better one, and you'll recompute anyway. What an open licence removes isn't the migration: it's having the date set by someone else.

Open weights don't mean exact reproducibility either. Change the runtime, the quantisation or the tokeniser version and vectors can shift slightly. For an index you've already built, sometimes that's fine and sometimes it isn't — measure it with your own queries rather than assuming.

And somebody has to keep those weights. If your plan is to download them from a public repository four years from now, your plan is that the repository is still there. Download them now and put them in your backups; they weigh less than the index does.

One last bit of honesty: it isn't a universal disaster either. Recomputing a mid-sized corpus with a small model is a matter of hours and not much money. This hurts at scale, or when you didn't keep the text.

Our take

I agree with Simon on the substance and I'd move the emphasis. The licence isn't your insurance: your insurance is being able to rebuild the index from the original documents on any given Tuesday, with whichever model. Apache 2.0 is the second line of defence, and a good one — but I've seen far more projects stuck on chunking nobody documented than on a retired model.

The rule I apply is simple: if you persist it, demand the ability to regenerate it; if it's ephemeral, use whatever you like. For chat I use whichever model is best this month and I don't care if it's closed. For anything that stays in the database, I want an exit that doesn't depend on somebody else's commercial roadmap.

I also think part of this debate will age oddly. Embedding models are small and getting cheaper to run; in two years, recomputing a whole index will be a weekend cron job and the lock-in will have moved somewhere else — the vector database, the graph, whatever you built on top. What won't age is the obligation to keep the source. That's been true of every search migration of the last twenty years and it'll stay true.

When someone asks us to set up semantic search over their internal documentation, the first question is never which model: it's where the texts are and whether they can be chunked exactly the same way again. Half the time, that's where the conversation about models ends and the real one begins.

Any questions, tell me and we'll go through it. All the best, Vicente.

Source: Simon Willison

Did reading this raise a question?

Ask us. We answer even if you never become a client.