PROJECT
JULY 2026 A video accessibility tool that generates frame-by-frame visual descriptions, scene-aware narration, and structured transcripts for visually impaired users.
Roles
AI Engineer
- Built a video accessibility platform using Python and Streamlit
- Generated scene-aware narration and structured transcripts for YouTube videos
- Developed a multimodal analysis pipeline integrating visual, audio, and contextual models
Technical Highlights
- Integrated Qwen2.5-VL-3B-Instruct, OpenAI Whisper, TransNetV2, CLIP ViT-B/32, and Qwen2.5 7B via Ollama
- Used TransNetV2 for scene change detection and CLIP for keyframe selection
- Applied Retrieval-Augmented Generation to refine descriptions with visual and conversational context
- Minimized compute latency and API costs through local caching frameworks
- Achieved a 4.62/5.0 system quality score