DIABLo: a Deep Individual-Agnostic Binaural Localizer

In this project, we have developed and studied a deep neural network-based individual-agnostic general-purpose binaural localizer (BL) for sound sources located at arbitrary directions on the $4\pi$ sphere. Unlike binaural localization models trained with an HRIR catalog associated with a specific head and ear shape, an individual-agnostic model aims for the generalization over the individuality of HRIRs, and does not assume a-priori knowledge about the HRIRs which the sound wave is filtered through at recording time. The proposed model was evaluated via localization tests using public binaural room impulse responses (BRIRs) and binaural recording datasets and was found to deliver more robust and accurate localization in noisy and reverberant conditions and unknown recording-time HRIRs compared to BLs trained on a single subject’s HRIR catalog. The proposed model is also designed to support multiple or moving sources, and demonstrations for these scenarios are provided.

发言人详细信息

Shoken Kaneko received the B. Eng. degree in Electronic Engineering from the University of Tokyo in 2008 and the M. Sci. degree in Physics from Radboud University in 2010. He has been working as a Research Engineer at Yamaha Corporation from 2011 to 2019 in the field of spatial audio for virtual/augmented reality and numerical acoustic simulations. His major contribution at Yamaha includes the development of a binaural spatial audio technology based on statistical ear shape modeling branded as ViReal headphones^TM, which was licensed to a major video game company and was deployed in some million-seller video games which have sold more than 22 million copies in total by May 2020. Currently, he is a Ph.D. student in the Computer Science program at the University of Maryland, College Park. His research interests include scientific computing methods for numerical acoustic simulation and immersive audio technologies for virtual/augmented reality and next-generation telecommunication/telepresence.

专题：: Microsoft Research Talks
日期：: 2021年8月12日
演讲者：: Shoken Kaneko
所属机构：: University of Maryland, College Park

- Shoken Kaneko
  
  Ph.D. student
  
  University of Maryland, College Park
- Hannes Gamper
  
  Principal Researcher
研究领域
- Audio and Acoustics
研究院
- Microsoft Research Lab - Redmond
组
- Audio and Acoustics Research Group
项目
- Spatial Audio
下载
- DIABLo: Deep individual-agnostic binaural localizer

系列： Microsoft Research Talks

Decoding the Human Brain – A Neurosurgeon’s Experience
August 1, 2024
Speakers:

Pascal Zinn,

Ivan Tashev
Scalable and Efficient AI: From Supercomputers to Smartphones
June 29, 2023
Speakers:

Torsten Hoefler
Human-Centered AI: Ensuring Human Control While Increasing Automation
May 3, 2023
WiDS Career Panel: Gabriela de Queiroz, Juliet Hougland, & Samantha Sifleet
April 5, 2023
Speakers:

Gabriela de Queiroz,

Juliet Hougland,

Samantha Silfleet
Galea: The Bridge Between Mixed Reality and Neurotechnology
February 13, 2023
Speakers:

Eva Esteban,

Conor Russomanno
Current and Future Application of BCIs
February 1, 2023
Speakers:

Christoph Guger
Challenges in Evolving a Successful Database Product (SQL Server) to a Cloud Service (SQL Azure)
October 27, 2022
Speakers:

Hanuma Kodavalla,

Phil Bernstein
Improving text prediction accuracy using neurophysiology
September 30, 2022
Speakers:

Sophia Mehdizadeh
Tongue-Gesture Recognition in Head-Mounted Displays
August 11, 2022
Speakers:

Tan Gemicioglu
DIABLo: a Deep Individual-Agnostic Binaural Localizer
August 12, 2021
Speakers:

Shoken Kaneko
A Tale of Two Cities: Software Developers in Practice During the COVID-19 Pandemic
February 26, 2021
Speakers:

Denae Ford Robinson
Recent Efforts Towards Efficient And Scalable Neural Waveform Coding
September 29, 2020
Speakers:

Kai Zhen
Geometry-constrained Beamforming Network for end-to-end Farfield Sound Source Separation
September 24, 2020
Speakers:

Ali Aroudi
Audio-based Toxic Language Detection
August 13, 2020
Speakers:

Midia Yousefi
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
August 4, 2020
Speakers:

Paul Smolensky,

Sean Andrist
From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
July 29, 2020
Speakers:

Sujeeth Bharadwaj
Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
July 29, 2020
Speakers:

Monojit Choudhury
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
July 28, 2020
Speakers:

Paul Smolensky,

Sean Andrist
An Ethical Crisis in Computing?
March 3, 2020
Speakers:

Emre Kiciman,

Eric Horvitz
Towards Mainstream Brain-Computer Interfaces (BCIs)
February 27, 2020
Speakers:

Hannes Gamper
Underestimating the challenge of cognitive disabilities (and digital literacy). Directions to explore for current, next, and next-next generation UIs
November 25, 2019
Speakers:

Gregg Vanderheiden
'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project
November 18, 2019
Speakers:

Peter Clark
Checkpointing the Un-checkpointable: the Split-Process Approach for MPI and Formal Verification
November 15, 2019
Speakers:

Gene Cooperman
Learning Structured Models for Safe Robot Control
September 27, 2019
Speakers:

Ashish Kapoor
Non-linear Invariants for Control-Command Systems
September 6, 2019
Speakers:

Tahina Ramananandro