Research scientist at NVIDIA working on speech and multi-modal LLMs. Previously PhD student at MIT CSAIL and intern at @AIatMeta and @Apple MLR.people.csail.mit.edu/roudi/Joined December 2016
@adamcohenhillel Nice demo! Lip reading has a lot of potential after being personalized to a specific user. You could consider using audio-visual lip reading methods like Whisper-Flamingo, LLaMA-AVSR, etc... for better recognition when the lips aren't clearly visible.
How do AVSR models balance what they hear and what they see?
Introducing Dr. SHAP-AV, the first large-scale Shapley-based analysis of modality contributions in audio-visual speech recognition.
6 sota models 2 benchmarks 3 analyses
๐Project page: umbertocappellazzo.github.io/Dr-SHAP-AV/
๐งต๐
Do you really need audio to fine-tune your Audio LLM? ๐ค Answer below:
Introducing Omni-R1, a simple GRPO fineโtuning method for Qwen2.5โOmni on audio question answering. It sets new stateโofโtheโart accuracies on the MMAU benchmark for Audio LLMs.
arxiv.org/abs/2505.09439
๐๏ธ Another #MultimodalAI workshop we are organizingโthis one zeroes in on speech & language foundation models!
๐ Dive into #SpeechAI, audio, and language tech. Learn how to build foundation models and hear from both academia and industry experts.
๐Sep 4โ5, 2025 | @TTIC_Connect
๐ข Excited to announce our 2-day workshop on "Foundations of Speech and Audio Foundation Models" at TTI Chicago, happening September 4โ5!
๐ Info & registration: sites.google.com/view/speech-aiโฆ
๐ Poster submissions welcome!
Join us for talks, discussions, and community building!
Congrats to Edson for leading our Contrastive Audio-Visual Masked Autoencoders 2.0 Project (CAV-MAE Sync), accepted at #CVPR2025!
Check out Edson's thread for more details โฌ๏ธ
Link the MMAU leaderboard (Massive Multi-Task Audio Understanding and Reasoning Benchmark) - it should hopefully be updated soon with Omni-R1 sakshi113.github.io/mmau_homepage/โฆ
``Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities,'' George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish Mittal, Brian Kingsbury, David Haws, Edmilson Morais, Gakuto Kurata, Haโฆ ift.tt/QPsxkH2
80 Followers 628 FollowingBiotechnologist| Documentary filmmaker|Politics and social affairs Commentator|Service-oriented not profit-motivated|Rwandan Patriot, Pan-Africanist|RWANDA 1st
5 Followers 384 FollowingMTech Research student @iiscbangalore
Education brain rot maxxing
I will treat this space as my diary I'll be logging updates.
595K Followers 54K FollowingSan Francisco/Silicon Valley AI | Robots, holodecks, BCIs, analysis of new things | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
102 Followers 1K Followinglatent space mechanic | experimenter in deep learning and musical audio | Ableton Live enjoyer | prev. cs @stanford, @StanfordAILab
1K Followers 5K FollowingHead of Data and Digital Preservation at the BFI National Archive. Working on digital preservation, audiovisual archiving, using AI with care. Opinions my own.
486K Followers 1K FollowingML/AI research engineer. Ex stats professor.
Author of "Build a Large Language Model From Scratch" (https://t.co/O8LAAMRzzW) & reasoning (https://t.co/5TueQKx2Fk)
96K Followers 927 FollowingOpen model research @ something new.
Prev. co-led Olmo at Ai2.
Contact via email.
Writes @interconnectsai
Wrote The RLHF Book,
๐๏ธ๐โโ๏ธ
112 Followers 134 FollowingI am an associate professor in the Department of Computer Science and Engineering at The Ohio State University. I lead The ASPIRE Group.
102 Followers 1K Followinglatent space mechanic | experimenter in deep learning and musical audio | Ableton Live enjoyer | prev. cs @stanford, @StanfordAILab
66 Followers 423 FollowingTalent Lead at Liquid AIโฆbuilding the team who are building LFMs. Donโt know what LFMs are? click here: https://t.co/M4pd4P14dy
123 Followers 147 FollowingPh.D. student at the Chinese University of Hong Kong, Shenzhen. Amphion co-founder. Interested in speech, music, and multi-modality.
50 Followers 2 FollowingIEEE ASRU, is focused on bringing together academia and industry to discuss new developments on the field of Automatic Speech Recognition and Understanding.
230K Followers 812 FollowingCooking fun AI systems & products @databricks. Prev: co-founder & CTO @ Hyperbolic, OctoAI (acquired by @nvidia) Apache TVM, PhD @ University of Washington.