← all posts
// analysis · embeddings

After word2vec: what Mikolov built next

word2vec is the most-cited thing Tomáš Mikolov ever put his name on, and almost nobody runs it in production anymore. The 2013 paper collected its NeurIPS Test of Time award in 2023, sailed past 40,000 citations, and Mikolov himself is now back in Prague working on something he's decided matters more. So here's a question people skip while they quote king minus man plus woman equals queen: what did the person who built word2vec build next? The answer runs in both directions from 2013, and the tool you'd still install today isn't the one on the citation.

Before word2vec, there was RNNLM

The part everyone forgets is 2010, three years earlier. Mikolov, finishing his PhD at Brno with Lukáš Burget and Martin Karafiát, plus Sanjeev Khudanpur over at Johns Hopkins, put a recurrent neural network on plain language modeling and called the toolkit RNNLM. It roughly halved perplexity against a strong backoff n-gram and cut about 18% off word error rate on Wall Street Journal. He's argued, not quietly, that this was as revolutionary as AlexNet: gradient clipping, neural text generation, and character-level modeling were all sitting in that toolkit by 2010. Buy the framing or don't, the point holds. The recurrent language model came first, and word2vec fell out of trying to make its input layer cheap.

word2vec: the paper that almost didn't get in

You know the famous half. Two 2013 papers, skip-gram and CBOW, word vectors where analogies drop out of plain arithmetic. The less famous half: the first one was rejected at the inaugural ICLR in 2013, a conference that accepted around 70% of what it received that year. What made word2vec win wasn't the math, it was shipping. A fast open C implementation anyone could run over a weekend on a laptop. Mikolov has since said the code only looked obfuscated because he'd over-optimized it while waiting for Google to clear the release. GloVe arrived a year later, borrowed a pile of the tricks, ran slower and heavier, and got popular anyway on the strength of bigger pretrained vectors. The better-marketed follow-up: remember that pattern, it comes back.

fastText is the successor you'd still pip install

If word2vec has a direct heir from Mikolov's own hand, it's fastText, 2016, out of Facebook AI Research with Piotr Bojanowski, Edouard Grave and Armand Joulin. One small idea, and it's exactly why the tool aged well: stop treating a word as an atom. fastText chops each word into character n-grams, three to six characters long, and sums their vectors, so a word is built out of its pieces rather than looked up whole.

  • Out-of-vocabulary words get a real vector instead of a zero, because their n-grams showed up inside other words during training.
  • Morphology comes for free. Prefixes, suffixes, roots and compounds land near each other, which matters enormously for Czech, Finnish or Turkish, exactly the languages word2vec handled worst.
  • Facebook released pretrained vectors for 157 languages, which is why fastText is still the default when someone needs static embeddings in a language that isn't English.

That's the practical nugget hiding in the history. When you genuinely need cheap, static, per-word vectors today, fastText is what you reach for. word2vec is what you cite.

word2vec is the paper you cite; fastText is the file you download. They're almost never the same object.

The successor that wasn't his

Now the honest part, because the lineage isn't a straight victory lap. The thing that knocked word2vec out of practice didn't come from Mikolov at all. ELMo and then BERT, both 2018, made embeddings contextual: the vector for bank now depends on the sentence wrapped around it, and that killed static embeddings for serious NLP more or less overnight. fastText is still excellent at its job, the job just got smaller. Doing classification or retrieval on real text in 2026, you're reaching for a transformer embedding model, not a bag of character n-grams. Choosing an embedding model walks that decision in detail; the compressed version is that the fixed vector lost to the contextual one, and it wasn't close.

What Mikolov is doing now

He left Facebook in 2020 and came home to Prague, to CIIRC at the Czech Technical University and the RICAIP centre. The telling thing is what he picked to work on, which is emphatically not bigger language models. He's openly skeptical that scaling LLMs on their own reaches general intelligence, and his group is after something weirder: complex systems that grow novel behavior on their own, closer to how life got complicated over evolutionary time, funded partly by a GoodAI grant aimed at precisely that. He gave an AGI-conference keynote titled AGI: Why and How, and picked up the Czech AI Award in 2024 on top of the NeurIPS honor. If you want the man rather than the CV, read his blunt post from around the award: still visibly annoyed that Quoc Le and Ilya Sutskever turned his end-to-end translation idea into sequence-to-sequence without listing him, still convinced RNNLM never got the credit AlexNet did. There's a whole argument about how far we are from AGI tucked inside that bet.

The tidy version of this story is that word2vec begat everything. The real one is messier and more useful. Mikolov's actual line runs RNNLM to word2vec to fastText, then swerves out of the embedding game entirely toward a problem he thinks the field is busy ignoring. Take one thing from it into your own stack: the paper a field canonizes and the tool that survives contact with production are rarely the same artifact. Cite word2vec if you like. Install fastText. And keep half an eye on whatever the guy who built both decides is worth his time, because his hit rate on the unglamorous-but-right problem is better than almost anyone's in the room.

#embeddings#nlp#fasttext#history