Skip to content

All Benchmarks

UCF101-AD

A benchmark built to evaluate MLLM "capacity for denial," a measurement of model agreeableness and the tendency to hallucinate actions occurring.

Background

Multimodal large language models (MLLMs) struggle to recognize when an activity does not happen despite strong contextual clues in video. This is a critical blind spot in modern video understanding, as real video understanding requires the ability to reason whether an action actually happened.

Methodology

Two clips, one where the action occurs and a contextually similar one where the action does not occur, are paired together. This method ensures that the model is actually verifying the motion that defines the action rather than relying on contextual cues such as background or object presence.

Dataset Structure

11,283 videos built on the UCF101 action categories. The Action-Denial split includes 7,059 negative clips for training and 3,549 for testing. The Action-Presence split consists of a test-only set of 675 clips.

Benchmark Measured Capabilities

Task Examples

  • Question

    Which choice accurately describes this video?

    Question type
    Multiple choice
    Options
    • A.A person or group of people is swimming breaststroke.
    • B.A person or group of people is swinging a golf club.
    • C.A person or group of people is juggling balls.
    • D.A person or group of people is engaged in the sport of fencing, using specialized swords, and following formal fencing rules and techniques
    • E.A person or group of people is playing basketball shooting at the basket.
    • F.A person or group of people is swinging a tennis racket and hitting the ball.
    • G.None of the choices accurately describes this video.
    • H.A person or group of people is spiking a volleyball causing it to hit the ground.
    • I.A person or group of people is bowling a ball down an alley knocking down the pins.
    • J.A person or group of people is swinging on a playground swing.
    • K.A person or group of people is playing billiards hitting shot with stick.
    Correct answer
    K. A person or group of people is playing billiards hitting shot with stick.
    Capability
    Action Understanding
  • Question

    Which choice accurately describes this video?

    Question type
    Multiple choice
    Options
    • A.A person or group of people is diving into water.
    • B.A person or group of people is playing billiards hitting shot with stick.
    • C.A person or group of people is playing a dhol drum.
    • D.A person or group of people is walking on hands in a handstand.
    • E.A person or group of people is spinning while dancing salsa.
    • F.A person or group of people is riding a horse.
    • G.A person or group of people is cutting hair.
    • H.A person or group of people is pole vaulting.
    • I.A person or group of people is swimming breaststroke.
    • J.A person or group of people is skydiving, free-falling straight downwards risking hitting the ground if there was no parachute.
    • K.None of the choices accurately describes this video.
    Correct answer
    K. None of the choices accurately describes this video.
    Capability
    Action Understanding

@article{abdullah2026learningtodeny,
title   = {Learning to Deny: Action Denial in Multimodal Large Language Models},
author  = {Raiyaan Abdullah and Shehreen Azad and Yogesh Singh Rawat},
journal = {arXiv preprint arXiv:2606.31187},
year    = {2026},
url     = {https://arxiv.org/abs/2606.31187},
}