Grin logo
de en es fr
Shop
GRIN Website
Publish your texts - enjoy our full service for authors
Go to shop › Computer Sciences - Artificial Intelligence

Multimodal Framework for Video Summarization Using Transformers and Deep Learning

Integrating Visual, Audio, and Textual Data

Title: Multimodal Framework for Video Summarization Using Transformers and Deep Learning

Research Paper (postgraduate) , 2026 , 58 Pages

Autor:in: Suryakanthi Tangirala (Author), Jinaga Neeraja (Author)

Computer Sciences - Artificial Intelligence
Excerpt & Details   Look inside the ebook
Summary Details

This text aims to develop a multimodal deep learning framework for automatically generating concise and meaningful summaries of lengthy video content. The proposed approach combines visual, audio, and textual information to identify relevant video segments while preserving essential context and reducing redundancy.

The framework integrates pretrained models such as Vision Transformers (ViT) for visual feature extraction, Whisper for speech recognition, and BART or DistilBERT for textual and contextual processing. Transformer-based architectures combine these representations to assess segment relevance and generate coherent summaries. The approach can support applications in education, entertainment, surveillance, news analysis, sports, and multimedia information retrieval.

Details

Title
Multimodal Framework for Video Summarization Using Transformers and Deep Learning
Subtitle
Integrating Visual, Audio, and Textual Data
Authors
Suryakanthi Tangirala (Author), Jinaga Neeraja (Author)
Publication Year
2026
Pages
58
Catalog Number
V1759421
ISBN (PDF)
9783389206256
ISBN (Book)
9783389206263
Language
English
Tags
Multimodal Video Summarization Deep Learning Transformers Vision Transformer Automatic Speech Recognition
Product Safety
GRIN Publishing GmbH
Quote paper
Suryakanthi Tangirala (Author), Jinaga Neeraja (Author), 2026, Multimodal Framework for Video Summarization Using Transformers and Deep Learning, Munich, GRIN Verlag, https://www.grin.com/document/1759421
Look inside the ebook
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
Excerpt from  58  pages
Grin logo
  • Grin.com
  • Shipping
  • Contact
  • Privacy
  • Terms
  • Imprint
  • Withdraw Contract