EventsThe 1st International Online Conference on Behavioral Sciences
Published
This submission belongs to the session S9. Experimental and Clinical Neurosciences of the event The 1st International Online Conference on Behavioral Sciences
Published date
27 Mar, 2026
Academic Editor
author-avatarJerrell Cassady
Citation
Bin Hu, Shahryar Wasif, Mitigating Temporal Confabulation and Improving Calibrated Perception in Real-Time Vision–Language Models, in Proceedings of The 1st International Online Conference on Behavioral Sciences, 1 April–3 April 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Mitigating Temporal Confabulation and Improving Calibrated Perception in Real-Time Vision–Language Models

1. Canadian Open Digital Health (OpenDH) program and Department of Clinical Neurosciences, Cumming school of Medicine, University of Calgary, Calgary T2N 1N4, Canada, Canada
Abstract

Introduction: Real-time vision–language models (VLMs) can exhibit “cognitive-like” failure patterns, including temporally unstable judgments, persistence of incorrect hypotheses, and overconfident confabulations under uncertainty. Conventional single-image benchmarks do not isolate these time-dependent behaviors. We present CMC (Confabulation Mitigation & Calibration), a lightweight reliability wrapper, and evaluate it within a shared cognitive testing platform that probes interpretable perceptual and metacognitive functions in both humans and AI.

Methods: The platform targets three functions relevant to hallucination-like errors: visual perception and orientation discrimination, temporal stability of belief across successive observations, and confidence calibration under time constraints. We used a Tumbling-E orientation task with randomized staircase difficulty to stress perceptual decision-making while recording response time, timeouts/abstentions, and step-to-step consistency. CMC combines selective re-analysis via change detection, a risk score that integrates temporal instability signals with uncertainty features, and a confirmation stage that routes high-risk outputs to verification or calibrated responses (verify/hedge/abstain). Human trials used a 3-second response budget; AI trials used a longer budget to distinguish perceptual failure from system latency.

Results: In 40 Tumbling-E trials, baseline accuracy was 50% (20/40) without CMC and improved to 75% (30/40) with CMC v1. CMC further reduced temporally unstable behaviors by escalating high-risk steps to verification and suppressing overconfident outputs when evidence was weak or timeouts were likely.

Conclusions: Evaluated as a cognitive testing problem rather than a static captioning task, CMC improves both accuracy and cognitive-faithful reliability, supporting rigorous neuroscience-aligned research and safer deployment of real-time VLM systems.

Keywords
Vision-language models
temporal hallucination
confabulation mitigation
confidence calibration
Benchmark of Subtyping Pathological Stimulus Persistence and Confabulation in Multimodal AI
Enhancing Cognitive Function and Reducing Fatigue Through Occupational Therapy: Evidence from Postoperative Care in Older Adults