This text aims to develop a multimodal deep learning framework for automatically generating concise and meaningful summaries of lengthy video content. The proposed approach combines visual, audio, and textual information to identify relevant video segments while preserving essential context and reducing redundancy.
The framework integrates pretrained models such as Vision Transformers (ViT) for visual feature extraction, Whisper for speech recognition, and BART or DistilBERT for textual and contextual processing. Transformer-based architectures combine these representations to assess segment relevance and generate coherent summaries. The approach can support applications in education, entertainment, surveillance, news analysis, sports, and multimedia information retrieval.
- Quote paper
- Suryakanthi Tangirala (Author), Jinaga Neeraja (Author), 2026, Multimodal Framework for Video Summarization Using Transformers and Deep Learning, Munich, GRIN Verlag, https://www.grin.com/document/1759421