NLM DIR Seminar Schedule
UPCOMING SEMINARS
RECENT SEMINARS
-
June 30, 2026 Jaya Srivastava
Disrupted Regulation of Essential Genes Mediates Dementias and Age-Associated Disorders -
June 11, 2026 Angela Jiang
Identification and Evolutionary Analysis of Steroid-Metabolism Enzymes in Gut Microbes -
June 10, 2026 Luda Diatchenko
New Insights on Pain Biology from Human Transcriptomics: How Stimulation of Immune Response Shapes Pain Resolution -
June 9, 2026 Pascal Mutz
Characterization of covalently closed circular RNA replicators detected in (meta)transcriptomic data -
June 4, 2026 Madeleine Clore
Explaining why AlphaFold struggles to predict mutational effects
Scheduled Seminars on May 12, 2022
Contact NLMDIRSeminarScheduling@mail.nih.gov with questions about this seminar.
Abstract:
Videos that semantically correspond to a text query provide highly condensed information that can give a complete answer to the query. Videos relevant to medical instructional questions (e.g., how to use a tourniquet) are especially useful for first aid, medical emergency, and education questions. However, the number of publicly available, benchmark datasets with medical instructional videos is nonexistent. Thus we introduce two new datasets to push research toward designing and comparing algorithms that can recognize medical instructional videos and locate visual answers from them to natural language queries. We propose the datasets, MedVidCL and MedVidQA, for the tasks of Medical Video Classification (MVC) and Medical Visual Answer Localization (MVAL), two tasks that emphasize multi-modal (language and video) understanding. The MedVidCL dataset includes 6117 annotated videos for the MVC task, while the MedVidQA dataset contains 3010 annotated questions with corresponding answer segments from 899 videos for the MVAL task. We have benchmarked both tasks with both datasets via deep learning models that set competitive and comparative baselines for future research.