GulliBench is a benchmark that purports to measure the "gullibility" of AI models.
"Current AI models are trained to be extremely good at solving hard problems (i.e., being very smart). But these problems are usually well-defined and have clear solutions. They are created in sterile environments, where data is standardized and things are mostly deterministic. In short, models are trained with the heavy assumption that the data, tools, and knowledge they are handed are undeniably pristine."
"As most of us know, that is not the case in real life. Data is messy and sometimes plain wrong. Tools are buggy and give unreliable results. Assumptions need to be revisited and rewritten."
"GulliBench probes a single, specific failure: taking the data at face value instead of reconciling it against the primary source it should agree with. That's one slice of a much larger category of gullibility failures. Models can be gullible in plenty of other ways: believing a buggy tool's output, accepting a false premise baked into the prompt, deferring to a confident-but-wrong user, following a planted instruction from a document. We don't touch any of that here. We think data-trust is a clean, measurable place to start, but definitely not the whole story."
Here's there most "gullible" top 10:
- Opus 5 - 49
- Fable 5 - 48
- Muse Spark 1.2 - 42
- Gemini 3.1 Pro - 23
- Kimi K3 - 18
- Grok 4.6 - 16
- Opus 4.8 - 16
- DeepSeek V4 Flash - 16
- GLM 5.2 - 15
- DeepSeek V4 Pro - 11
GulliBench: Intelligence isn't enough. Measuring skepticism in frontier models.
#solidstatelife #ai #genai #llms #codingai #datascience
GulliBench: Intelligence Isn't Enough. Measuring Skepticism in Frontier Models.
GulliBench measures AI gullibility: on tiny, easy tasks, a convenient stored field lies while the truth sits one derivation away in the same data. Does the model check, or grab the number it was handed?Arthur Kamienski (Vetto)