3,266 questions about wine — that's how many made it into the new OenoBench benchmark. Its author spent six months collecting 38,104 facts from open sources and turned them into a multiple-choice test. The goal is to measure how much language models actually know about wine. In his report, he shares exactly where the models let him down. The benchmark has already been released, and you can try it on your own LLM. It's a good way to test a model on a niche topic or simply find out how well AI knows its wine.
Source: habr.com · post in Telegram