Skip to main content
Visonic AI Visonic AI
Research / Technical Visonic AD 2
Technical report

Long-form video understanding for audio description

Visonic AD 2 tracks context across a complete title, selects the visual information needed to follow it and places description inside the available soundtrack gaps.

Context across a complete title

Processing a title as isolated clips loses connections between distant scenes. A glance in minute forty may only make sense because of a decision in minute three. A returning location, costume or prop may carry meaning that is invisible when a scene is processed alone.

AD 2 carries context across the title and selects the visual information needed to follow the story.

Five linked decisions

The system has to recognise what is on screen, maintain people and relationships across time, determine narrative relevance, locate usable gaps in the soundtrack and write a description that fits those gaps.

An error at any one layer is visible in the final script. Accurate object recognition can still produce poor AD if it interrupts dialogue, repeats audible information or misses the action that changes the scene.

Why AD 2 has three models

A feature film, an episodic catalogue and a legacy archive have different production budgets and review requirements. Ultra serves flagship productions, Prime serves daily production and Micro serves larger libraries.

Improving the first result

Since the first Visonic Audio Description release in September 2023, we have tested model output against difficult titles and feedback from AD professionals. We use those cases to improve continuity, select more relevant action and reduce the edits needed before delivery.

Review requirements

Audio description contains editorial judgement. Ambiguous action, dense dialogue, specialised terminology, sensitive representation and exact house style can still require human review.