Stop the war!
Остановите войну!
for scientists:
default search action
IEEE/ACM Transactions on Audio, Speech and Language Processing, Volume 32
Volume 32, 2024
- Jin Chu Wu, Raghu N. Kacker:
Statistical Analysis for Speaker Recognition Evaluation With Data Dependence and Three Score Distributions. 1-14 - Yongwei Zhou, Junwei Bao, Youzheng Wu, Xiaodong He, Tiejun Zhao:
Operation-Augmented Numerical Reasoning for Question Answering. 15-28 - Anurenjan Purushothaman, Debottam Dutta, Rohit Kumar, Sriram Ganapathy:
Speech Dereverberation With Frequency Domain Autoregressive Modeling. 29-38 - Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, Stefan Wermter:
Disentangling Prosody Representations With Unsupervised Speech Reconstruction. 39-54 - Mathias Bach Pedersen, Søren Holdt Jensen, Zheng-Hua Tan, Jesper Jensen:
Data-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence Probability. 55-67 - Yuanbo Hou, Bo Kang, Andrew Mitchell, Wenwu Wang, Jian Kang, Dick Botteldooren:
Cooperative Scene-Event Modelling for Acoustic Scene Classification. 68-82 - Xiaotong Jiang, Peiwen You, Chen Chen, Zhongqing Wang, Guodong Zhou:
Exploring Scope Detection for Aspect-Based Sentiment Analysis. 83-94 - Xuenan Xu, Zeyu Xie, Mengyue Wu, Kai Yu:
Beyond the Status Quo: A Contemporary Survey of Advances and Challenges in Audio Captioning. 95-112 - Federico Miotello, Mirco Pezzoli, Luca Comanducci, Fabio Antonacci, Augusto Sarti:
Deep Prior-Based Audio Inpainting Using Multi-Resolution Harmonic Convolutional Neural Networks. 113-123 - Cristian Lucian Stanciu, Jacob Benesty, Constantin Paleologu, Ruxandra-Liana Costea, Laura-Maria Dogariu, Silviu Ciochina:
Decomposition-Based Wiener Filter Using the Kronecker Product and Conjugate Gradient Method. 124-138 - Huiyao Chen, Yueheng Sun, Meishan Zhang, Min Zhang:
Automatic Noise Generation and Reduction for Text Classification. 139-150 - Jiaming Xu, Jian Cui, Yunzhe Hao, Bo Xu:
Multi-Cue Guided Semi-Supervised Learning Toward Target Speaker Separation in Real Environments. 151-163 - Yang Xiang, Jesper Lisby Højvang, Morten Højfeldt Rasmussen, Mads Græsbøll Christensen:
A Two-Stage Deep Representation Learning-Based Speech Enhancement Method Using Variational Autoencoder and Adversarial Training. 164-177 - Xiao Li, Ruirui Liu, Huichou Huang, Qingyao Wu:
Contrastive Learning for Target Speaker Extraction With Attention-Based Fusion. 178-188 - Xiaobo Liang, Runze Mao, Lijun Wu, Juntao Li, Min Zhang, Qing Li:
Enhancing Low-Resource NLP by Consistency Training With Data and Model Perturbations. 189-199 - Haisheng Lu, Jiangnan Liang, Chuang Shi:
Comments on "Primary-Ambient Extraction Using Ambient Spectrum Estimation for Immersive Spatial Audio Reproduction". 200-202 - Szymon Drgas, Lars Bramsløw, Archontis Politis, Gaurav Naithani, Tuomas Virtanen:
Dynamic Processing Neural Network Architecture for Hearing Loss Compensation. 203-214 - Femke B. Gelderblom, Tron V. Tronstad, Torbjørn Svendsen, Tor André Myrvoll:
On the Predictive Power of Objective Intelligibility Metrics for the Subjective Performance of Deep Complex Convolutional Recurrent Speech Enhancement Networks. 215-226 - Thomas Haubner, Andreas Brendel, Walter Kellermann:
End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation. 227-238 - Congcong Jiang, Tieyun Qian, Bing Liu:
One General Teacher for Multi-Data Multi-Task: A New Knowledge Distillation Framework for Discourse Relation Analysis. 239-249 - Khandokar Md. Nayem, Donald S. Williamson:
Attention-Based Speech Enhancement Using Human Quality Perception Modeling. 250-260 - Ying Zhang, Fandong Meng, Yufeng Chen, Jinan Xu, Jie Zhou:
Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information Extraction. 261-275 - Jingsong Yan, Piji Li, Haibin Chen, Junhao Zheng, Qianli Ma:
Does the Order Matter? A Random Generative Way to Learn Label Hierarchy for Hierarchical Text Classification. 276-285 - Georgios Paraskevopoulos, Theodoros Kouzelis, Georgios Rouvalis, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos:
Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems: A Case Study for Modern Greek. 286-299 - Ernesto Accolti, Javier Gimenez, Michael Vorländer:
Uncertainties of Room Acoustics Simulation Due to Directivity Data of Musical Instruments. 300-309 - Yoshiki Masuyama, Kouei Yamaoka, Yuma Kinoshita, Taishi Nakashima, Nobutaka Ono:
Causal and Relaxed-Distortionless Response Beamforming for Online Target Source Extraction. 310-324 - Rohit Prabhavalkar, Takaaki Hori, Tara N. Sainath, Ralf Schlüter, Shinji Watanabe:
End-to-End Speech Recognition: A Survey. 325-351 - Yun Zhao, Dexi Liu, Changxuan Wan, Xiping Liu, Jian-Yun Nie, Jiaming Liu:
JMS-QA: A Joint Hierarchical Architecture for Mental Health Question Answering. 352-363 - Shiwen Ni, Jiawen Li, Min Yang, Hung-Yu Kao:
DropAttack: A Random Dropped Weight Attack Adversarial Training for Natural Language Understanding. 364-373 - Tiantian Zhu, Yang Qin, Ming Feng, Qingcai Chen, Baotian Hu, Yang Xiang:
BioPRO: Context-Infused Prompt Learning for Biomedical Entity Linking. 374-385 - Jiapu Wang, Boyue Wang, Junbin Gao, Simin Hu, Yongli Hu, Baocai Yin:
Multi-Level Interaction Based Knowledge Graph Completion. 386-396 - Qiangqiang Zhang, Dongyuan Lin, Yingying Xiao, Yunfei Zheng, Shiyuan Wang:
Error Reused Filtered-X Least Mean Square Algorithm for Active Noise Control. 397-412 - Zengrui Jin, Mengzhe Geng, Jiajun Deng, Tianzi Wang, Shujie Hu, Guinan Li, Xunying Liu:
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition. 413-429 - Jun Kong, Jin Wang, Xuejie Zhang:
Adaptive Ensemble Self-Distillation With Consistent Gradients for Fast Inference of Pretrained Language Models. 430-442 - Srdan Kitic, Jérôme Daniel:
Blind Identification of Ambisonic Reduced Room Impulse Response. 443-458 - Qijie Shao, Pengcheng Guo, Jinghao Yan, Pengfei Hu, Lei Xie:
Decoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent Recognition. 459-470 - Han Zhu, Gaofeng Cheng, Jindong Wang, Wenxin Hou, Pengyuan Zhang, Yonghong Yan:
Boosting Cross-Domain Speech Recognition With Self-Supervision. 471-485 - Yile Wang, Yue Zhang, Peng Li, Yang Liu:
Gradual Syntactic Label Replacement for Language Model Pre-Training. 486-496 - Penghui Ma, Jianfeng Li, Jingjing Pan, Xiaofei Zhang, Roberto Gil-Pita:
Coherent Signal DOA Estimation With Coprime Array: Exploiting Signal Subspace Reconstructing Strategy. 497-508 - Emma Hamel, Nickvash Kani:
Factors That Influence Automatic Recognition of African-American Vernacular English in Machine-Learning Models. 509-516 - Jingbei Li, Sipan Li, Ping Chen, Luwen Zhang, Yi Meng, Zhiyong Wu, Helen Meng, Qiao Tian, Yuping Wang, Yuxuan Wang:
Joint Multiscale Cross-Lingual Speaking Style Transfer With Bidirectional Attention Mechanism for Automatic Dubbing. 517-528 - Bing Han, Zhengyang Chen, Yanmin Qian:
Self-Supervised Learning With Cluster-Aware-DINO for High-Performance Robust Speaker Verification. 529-541 - Kristina Tesch, Timo Gerkmann:
Multi-Channel Speech Separation Using Spatially Selective Deep Non-Linear Filters. 542-553 - Hao-Chen Pei, Hao Fang, Xin Luo, Xin-Shun Xu:
Gradformer: A Framework for Multi-Aspect Multi-Granularity Pronunciation Assessment. 554-563 - Garima Sharma, Karthikeyan Umapathy, Sridhar Krishnan:
Time-Frequency Scattergrams for Biomedical Audio Signal Representation and Classification. 564-576 - Zhibo Man, Zengcheng Huang, Yujie Zhang, Yu Li, Yuanmeng Chen, Yufeng Chen, Jinan Xu:
WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation Learning. 577-590 - Chin-Po Chen, Ho-Hsien Pan, Susan Shur-Fen Gau, Chi-Chun Lee:
Using Measures of Vowel Space for Autistic Traits Characterization. 591-607 - Kevin Wilkinghoff, Frank Kurth:
Why Do Angular Margin Losses Work Well for Semi-Supervised Anomalous Sound Detection? 608-622 - Aku Rouhe, Tamás Grósz, Mikko Kurimo:
Principled Comparisons for End-to-End Speech Recognition: Attention vs Hybrid at the 1000-Hour Scale. 623-638 - Yile Wang, Yue Zhang:
Lost in Context? On the Sense-Wise Variance of Contextualized Word Embeddings. 639-650 - Christoph Hold, Ville Pulkki, Archontis Politis, Leo McCormack:
Compression of Higher-Order Ambisonic Signals Using Directional Audio Coding. 651-665 - Shouhui Wang, Biao Qin:
A Novel Joint Training Model for Knowledge Base Question Answering. 666-679 - Songbin Li, Jingang Wang, Peng Liu, Ke Shi:
SANet: A Compressed Speech Encoder and Steganography Algorithm Independent Steganalysis Deep Neural Network. 680-690 - Tarek Kanan, Amani AbedAlghafer, Shadi AlZu'bi, Bilal Hawashin, Ala Mughaid, Ghassan Kanaan, M. M. Kamruzzaman:
An Intelligent Health Care System for Detecting Drug Abuse in Social Media Platforms Based on Low Resource Language. 691-703 - Alejandro Santorum Varela, Svetlana Stoyanchev, Simon Keizer, Rama Doddipatla, Kate Knill:
Entity Resolution in Situated Dialog With Unimodal and Multimodal Transformers. 704-713 - Huang He, Hua Lu, Siqi Bao, Fan Wang, Hua Wu, Zheng-Yu Niu, Haifeng Wang:
Learning to Select External Knowledge With Multi-Scale Negative Sampling. 714-720 - Hua Lu, Zhen Guo, Chanjuan Li, Yunyi Yang, Huang He, Siqi Bao:
Towards Building an Open-Domain Dialogue System Incorporated With Internet Memes. 721-726 - Jungwoo Lim, Taesun Whang, Dongyub Lee, Heuiseok Lim:
Adaptive Multi-Domain Dialogue State Tracking on Spoken Conversations. 727-732 - David Thulke, Nico Daheim, Christian Dugast, Hermann Ney:
Task-Oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC10. 733-741 - Han Wu, Kun Xu, Linqi Song:
Structure-Aware Dialogue Modeling Methods for Conversational Semantic Role Labeling. 742-752 - Zhe Chen, Hongcheng Liu, Yu Wang:
DialogMCF: Multimodal Context Flow for Audio Visual Scene-Aware Dialog. 753-764 - Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang, Yang Feng, Jie Zhou, Seokhwan Kim, Yang Liu, Di Jin, Alexandros Papangelis, Karthik Gopalakrishnan, Dilek Hakkani-Tur, Babak Damavandi, Alborz Geramifard, Chiori Hori, Ankit Shah, Chen Zhang, Haizhou Li, João Sedoc, Luis F. D'Haro, Rafael E. Banchs, Alexander Rudnicky:
Overview of the Tenth Dialog System Technology Challenge: DSTC10. 765-778 - Shekhar Kumar Yadav, Nithin V. George:
Joint Dereverberation and Beamforming With Blind Estimation of the Shape Parameter of the Desired Source Prior. 779-793 - Yanxiong Li, Zhongjie Jiang, Qisheng Huang, Wenchang Cao, Jialong Li:
Lightweight Speaker Verification Using Transformation Module With Feature Partition and Fusion. 794-806 - Yuhan Dai, Zhirui Zhang, Yichao Du, Shengcai Liu, Lemao Liu, Tong Xu:
Datastore Distillation for Nearest Neighbor Machine Translation. 807-817 - Changtao Li, Feiran Yang, Jun Yang:
A Two-Stage Approach to Quality Restoration of Bone-Conducted Speech. 818-829 - Jie Zhou, Yuanbiao Lin, Qin Chen, Qi Zhang, Xuanjing Huang, Liang He:
CausalABSC: Causal Inference for Aspect Debiasing in Aspect-Based Sentiment Classification. 830-840 - Ruiying Lu, Bo Chen, Dandan Guo, Dongsheng Wang, Mingyuan Zhou:
Hierarchical Topic-Aware Contextualized Transformers. 841-852 - Yaru Zhao, Bo Cheng, Yakun Huang, Zhiguo Wan:
FluGCF: A Fluent Dialogue Generation Model With Coherent Concept Entity Flow. 853-867 - Changhao Ding, Zhangjie Fu, Zhongliang Yang, Qi Yu, Daqiu Li, Yongfeng Huang:
Context-Aware Linguistic Steganography Model Based on Neural Machine Translation. 868-878 - Zainab Alhakeem, Se-In Jang, Hong-Goo Kang:
Disentangled Representations in Local-Global Contexts for Arabic Dialect Identification. 879-890 - Jae-Hong Lee, Joon-Hyuk Chang:
Partitioning Attention Weight: Mitigating Adverse Effect of Incorrect Pseudo-Labels for Self-Supervised ASR. 891-905 - Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura:
Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech Segmentation. 906-916 - Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso:
Selective Acoustic Feature Enhancement for Speech Emotion Recognition With Noisy Speech. 917-929 - Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu:
Spatially Selective Speaker Separation Using a DNN With a Location Dependent Feature Extraction. 930-945 - Matan Karo, Arie Yeredor, Itshak Lapidot:
Compact Time-Domain Representation for Logical Access Spoofed Audio. 946-958 - Or Berebi, Zamir Ben-Hur, David Lou Alon, Boaz Rafaely:
Analysis and Design of Head-Tracked Compensation for Bilateral Ambisonics. 959-972 - Wei Wang, Yanmin Qian:
Universal Cross-Lingual Data Generation for Low Resource ASR. 973-983 - Davide Berghi, Philip J. B. Jackson:
Leveraging Visual Supervision for Array-Based Active Speaker Detection and Localization. 984-995 - Daniel Aleksander Krause, Guillermo García-Barrios, Archontis Politis, Annamaria Mesaros:
Binaural Sound Source Distance Estimation and Localization for a Moving Listener. 996-1011 - Seung-Bin Kim, Sang-Hoon Lee, Ha-Yeong Choi, Seong-Whan Lee:
Audio Super-Resolution With Robust Speech Representation Learning of Masked Autoencoder. 1012-1022 - Omer Musa Battal, Aykut Koç:
Automatic Construction of Sememe Knowledge Bases From Machine Readable Dictionaries. 1023-1035 - Varun Krishna, Tarun Sai, Sriram Ganapathy:
Representation Learning With Hidden Unit Clustering for Low Resource Speech Applications. 1036-1047 - Zhengding Luo, Dongyuan Shi, Woon-Seng Gan, Qirui Huang:
Delayless Generative Fixed-Filter Active Noise Control Based on Deep Learning and Bayesian Filter. 1048-1060 - Zewen Chi, Heyan Huang, Luyang Liu, Yu Bai, Xiaoyan Gao, Xian-Ling Mao:
Can Pretrained English Language Models Benefit Non-English NLP Systems in Low-Resource Scenarios? 1061-1074 - Rui Liu, Yifan Hu, Haolin Zuo, Zhaojie Luo, Longbiao Wang, Guanglai Gao:
Text-to-Speech for Low-Resource Agglutinative Language With Morphology-Aware Language Model Pre-Training. 1075-1087 - Shu Jiang, Zuchao Li, Hai Zhao, Weiping Ding:
Entity-Relation Extraction as Full Shallow Semantic Dependency Parsing. 1088-1099 - Yoav Vered, Stephen J. Elliott:
A Parallel Analog and Digital Adaptive Feedforward Controller for Active Noise Control. 1100-1108 - Puning Zhang, Rongjian Zhao, Boran Yang, Yuexian Li, Zhigang Yang:
Integrated Syntactic and Semantic Tree for Targeted Sentiment Classification Using Dual-Channel Graph Convolutional Network. 1109-1124 - Xu Wang, Hainan Zhang, Shuai Zhao, Hongshen Chen, Zhuoye Ding, Zhiguo Wan, Bo Cheng, Yanyan Lan:
Debiasing Counterfactual Context With Causal Inference for Multi-Turn Dialogue Reasoning. 1125-1132 - Hoang Ngoc Chau, Tien Dat Bui, Huu Binh Nguyen, Thanh Thi Hien Duong, Quoc-Cuong Nguyen:
A Novel Approach to Multi-Channel Speech Enhancement Based on Graph Neural Networks. 1133-1144 - Yuchen Hu, Chen Chen, Qiushi Zhu, Eng Siong Chng:
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR. 1145-1156 - Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino:
Blind and Spatially-Regularized Online Joint Optimization of Source Separation, Dereverberation, and Noise Reduction. 1157-1172 - Vibhav Agarwal, Sourav Ghosh, Harichandana B. S. S, Himanshu Arora, Barath Raj Kandur Raja:
TrICy: Trigger-Guided Data-to-Text Generation With Intent Aware Attention-Copy. 1173-1184 - Christoph Böddeker, Aswin Shanmugam Subramanian, Gordon Wichern, Reinhold Haeb-Umbach, Jonathan Le Roux:
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings. 1185-1197 - Reza Varzandeh, Simon Doclo, Volker Hohmann:
Speech-Aware Binaural DOA Estimation Utilizing Periodicity and Spatial Features in Convolutional Neural Networks. 1198-1213 - Yigitcan Özer, Meinard Müller:
Source Separation of Piano Concertos Using Musically Motivated Augmentation Techniques. 1214-1225 - Lior Frenkel, Shlomo E. Chazan, Jacob Goldberger:
Domain Adaptation Using Suitable Pseudo Labels for Speech Enhancement and Dereverberation. 1226-1236 - Jiahao Zhao, Wenji Mao, Daniel Dajun Zeng:
Disentangled Text Representation Learning With Information-Theoretic Perspective for Adversarial Robustness. 1237-1247 - Dong Zhou, Fang Lei, Lin Li, Yongmei Zhou, Aimin Yang:
Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics Retrieval. 1248-1260 - Xuechen Liu, Md. Sahidullah, Kong Aik Lee, Tomi Kinnunen:
Generalizing Speaker Verification for Spoof Awareness in the Embedding Space. 1261-1273 - Shiyao Cui, Jiangxia Cao, Xin Cong, Jiawei Sheng, Quangang Li, Tingwen Liu, Jinqiao Shi:
Enhancing Multimodal Entity and Relation Extraction With Variational Information Bottleneck. 1274-1285 - Yizhou Tan, Haojun Ai, Shengchen Li, Mark D. Plumbley:
Acoustic Scene Classification Across Cities and Devices via Feature Disentanglement. 1286-1297 - Orel Ben Zaken, Anurag Kumar, Vladimir Tourbabin, Boaz Rafaely:
Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information. 1298-1309 - Changsheng Quan, Xiaofei Li:
SpatialNet: Extensively Learning Spatial Information for Multichannel Joint Speech Separation, Denoising and Dereverberation. 1310-1323 - Matthew Baas, Herman Kamper:
Disentanglement in a GAN for Unconditional Speech Synthesis. 1324-1335 - Xian Li, Nian Shao, Xiaofei Li:
Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks. 1336-1351 - Yifan Chen, Gaofeng Cheng, Runyan Yang, Pengyuan Zhang, Yonghong Yan:
Interrelate Training and Clustering for Online Speaker Diarization. 1352-1364 - Sheng Feng, Xiaoqian Zhu, Shuqing Ma:
Masking Hierarchical Tokens for Underwater Acoustic Target Recognition With Self-Supervised Learning. 1365-1379 - Yangyang Zhao, Kai Yin, Zhenyu Wang, Mehdi Dastani, Shihan Wang:
Decomposed Deep Q-Network for Coherent Task-Oriented Dialogue Policy Learning. 1380-1391 - Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Gaël Richard, Florence d'Alché-Buc:
Tackling Interpretability in Audio Classification Networks With Non-negative Matrix Factorization. 1392-1405 - Xiuying Chen, Shen Gao, Mingzhe Li, Qingqing Zhu, Xin Gao, Xiangliang Zhang:
Write Summary Step-by-Step: A Pilot Study of Stepwise Summarization. 1406-1415 - Changkai Lin, Hongju Cheng, Qiang Rao, Yang Yang:
M$^{3}$SA: Multimodal Sentiment Analysis Based on Multi-Scale Feature Extraction and Multi-Task Learning. 1416-1429 - Ritujoy Biswas, Karan Nathwani, Vinayak Abrol:
Statistically Guided Near-End Speech Intelligibility Improvement Through Voice Transformation and Transfer Learning. 1445-1456 - Linhui Sun, Shuo Yuan, Aifei Gong, Lei Ye, Eng Siong Chng:
Dual-Branch Modeling Based on State-Space Model for Speech Enhancement. 1457-1467 - Alkis Koudounas, Eliana Pastor, Giuseppe Attanasio, Vittorio Mazzia, Manuel Giollo, Thomas Gueudré, Elisa Reale, Luca Cagliero, Sandro Cumani, Luca de Alfaro, Elena Baralis, Daniele Amberti:
Towards Comprehensive Subgroup Performance Analysis in Speech Models. 1468-1480 - Wenmeng Xiong, Changchun Bao, Jing Zhou, Maoshen Jia, José Picheral:
Joint DOA Estimation and Dereverberation Based on Multi-Channel Linear Prediction Filtering and Azimuth Sparsity. 1481-1493 - Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling:
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement. 1430-1444 - Yehav Alkaher, Israel Cohen:
Howling Detection and Gain Control for Speech Reinforcement in a Noisy Car Cabin Environment. 1494-1505 - Xinfa Zhu, Yi Lei, Tao Li, Yongmao Zhang, Hongbin Zhou, Heng Lu, Lei Xie:
METTS: Multilingual Emotional Text-to-Speech by Cross-Speaker and Cross-Lingual Emotion Transfer. 1506-1518 - Myeonghun Jeong, Minchan Kim, Byoung Jin Choi, Jaesam Yoon, Won Jang, Nam Soo Kim:
Transfer Learning for Low-Resource, Multi-Lingual, and Zero-Shot Multi-Speaker Text-to-Speech. 1519-1530 - Jiadi Yao, Hong Luo, Jun Qi, Xiao-Lei Zhang:
Interpretable Spectrum Transformation Attacks to Speaker Recognition Systems. 1531-1545 - Xiang Chen, Lei Li, Yuqi Zhu, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, Ningyu Zhang, Huajun Chen:
Sequence Labeling as Non-Autoregressive Dual-Query Set Generation. 1546-1558 - Lei Liu, Li Liu, Haizhou Li:
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition. 1559-1572 - Adrián Barahona-Ríos, Tom Collins:
NoiseBandNet: Controllable Time-Varying Neural Synthesis of Sound Effects Using Filterbanks. 1573-1585 - Siyuan Wang, Zhongyu Wei, Jiarong Xu, Taishan Li, Zhihao Fan:
Unifying Structure Reasoning and Language Pre-Training for Complex Reasoning Tasks. 1586-1595 - Yijing Chu, Sipei Zhao, Feng Niu, Yongzheng Dong, Yuezhe Zhao:
A New Diffusion Filtered-X Affine Projection Algorithm: Performance Analysis and Application in Windy Environment. 1596-1608 - Yuquan Le, Zhe Quan, Jiawei Wang, Da Cao, Kenli Li:
$\boldsymbol{R}^{2}$: A Novel Recall & Ranking Framework for Legal Judgment Prediction. 1609-1622 - Xiaotong Jiang, Ruirui Bai, Zhongqing Wang, Guodong Zhou:
Cross-Domain Aspect-Based Sentiment Classification With Tripartite Graph Modeling. 1623-1635 - Zhengyang Chen, Bing Han, Shuai Wang, Yanmin Qian:
Attention-Based Encoder-Decoder End-to-End Neural Diarization With Embedding Enhancer. 1636-1649 - Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jun Ma, Shaojun Wang, Jing Xiao:
EfficientTTS 2: Variational End-to-End Text-to-Speech Synthesis and Voice Conversion. 1650-1661 - Orel Peretz, Israel Cohen:
Constant Elevation-Beamwidth Beamforming With Concentric Ring Arrays. 1662-1672 - Zhibin Quan, Chi-Man Vong, Weili Zeng, Wankou Yang:
The MorPhEMe Machine: An Addressable Neural Memory for Learning Knowledge-Regularized Deep Contextualized Chinese Embedding. 1673-1686 - Lijian Gao, Qirong Mao, Ming Dong:
On Local Temporal Embedding for Semi-Supervised Sound Event Detection. 1687-1698 - Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu, Haizhou Li:
Accented Text-to-Speech Synthesis With Limited Data. 1699-1711 - Vinay Kothapally, John H. L. Hansen:
Monaural Speech Dereverberation Using Deformable Convolutional Networks. 1712-1723 - Taihui Wang, Feiran Yang, Jun Yang:
Multichannel Linear Prediction-Based Speech Dereverberation Considering Sparse and Low-Rank Priors. 1724-1735 - Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez, Piotr Zelasko, Najim Dehak:
Time-Domain Speech Super-Resolution With GAN Based Modeling for Telephony Speaker Verification. 1736-1749 - Marco Olivieri, Amy Bastine, Mirco Pezzoli, Fabio Antonacci, Thushara D. Abhayapala, Augusto Sarti:
Acoustic Imaging With Circular Microphone Array: A New Approach for Sound Field Analysis. 1750-1761 - Tengfei Liu, Yongli Hu, Junbin Gao, Yanfeng Sun, Baocai Yin:
Hierarchical Multi-Granularity Interaction Graph Convolutional Network for Long Document Classification. 1762-1775 - Etienne Thuillier, Craig T. Jin, Vesa Välimäki:
HRTF Interpolation Using a Spherical Neural Process Meta-Learner. 1790-1802 - Xun Gong, Yu Wu, Jinyu Li, Shujie Liu, Rui Zhao, Xie Chen, Yanmin Qian:
Advanced Long-Content Speech Recognition With Factorized Neural Transducer. 1803-1815 - Yoshiki Masuyama, Kouei Yamaoka, Takao Kawamura, Nobutaka Ono:
Efficient Joint Optimization of Sampling Rate Offsets Using Entire Multichannel Signal. 1816-1828 - Takaaki Saeki, Soumi Maiti, Xinjian Li, Shinji Watanabe, Shinnosuke Takamichi, Hiroshi Saruwatari:
Text-Inductive Graphone-Based Language Adaptation for Low-Resource Speech Synthesis. 1829-1844 - Douglas D. O'Shaughnessy:
Review of Methods for Automatic Speaker Verification. 1776-1789 - Yingming Gao, Peter Birkholz, Ya Li:
Articulatory Copy Synthesis Based on the Speech Synthesizer VocalTractLab and Convolutional Recurrent Neural Networks. 1845-1858 - Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas:
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection. 1859-1872 - Luciana M. X. de Souza, Márcio H. Costa, Renata Coelho Borges:
Envelope-Based Multichannel Noise Reduction for Cochlear Implant Applications. 1873-1884 - Linjian Li, Yi Cai, Xin Wu:
Unsupervised Disentanglement Learning Model for Exemplar-Guided Paraphrase Generation. 1885-1900 - Amir Ivry, Israel Cohen, Baruch Berdugo:
A User-Centric Approach for Deep Residual-Echo Suppression in Double-Talk. 1901-1914 - Geng Zhang, Jin Liu, Guangyou Zhou, Kunsong Zhao, Zhiwen Xie, Bo Huang:
Question-Directed Reasoning With Relation-Aware Graph Attention Network for Complex Question Answering Over Knowledge Graph. 1915-1927 - Yu Yao, Peng Yang, Guangzhen Zhao, Guoshun Yin:
KGAgent: Learning a Deep Reinforced Agent for Keyphrase Generation. 1928-1940 - Jiahong Li, Chenda Li, Yifei Wu, Yanmin Qian:
Unified Cross-Modal Attention: Robust Audio-Visual Speech Recognition and Beyond. 1941-1953 - Mieszko Fras, Konrad Kowalczyk:
Reverberant Source Separation Using NTF With Delayed Subsources and Spatial Priors. 1954-1967 - Rui Wang, Li Li, Tomoki Toda:
Dual-Channel Target Speaker Extraction Based on Conditional Variational Autoencoder and Directional Information. 1968-1979 - Qinyu Han, Zhihao Yang, Hongfei Lin, Tian Qin:
Let Topic Flow: A Unified Topic-Guided Segment-Wise Dialogue Summarization Framework. 2021-2032 - Haonan Cheng, Shulin Liu, Zhicheng Lian, Long Ye, Qin Zhang:
MusicECAN: An Automatic Denoising Network for Music Recordings With Efficient Channel Attention. 2033-2049 - Guy Gubnitky, Roee Diamant:
Detecting the Presence of Sperm Whales' Echolocation Clicks in Noisy Environments. 2050-2061 - Yuxia Wu, Tianhao Dai, Zhedong Zheng, Lizi Liao:
Active Discovering New Slots for Task-Oriented Conversation. 2062-2072 - Jacob Hollebon, Filippo Maria Fazi:
Dynamic Higher-Order Stereophony. 2073-2084 - Aidan O. T. Hogg, Mads Jenkins, He Liu, Isaac Squires, Samuel J. Cooper, Lorenzo Picinali:
HRTF Upsampling With a Generative Adversarial Network Using a Gnomonic Equiangular Projection. 2085-2099 - Yusheng Liao, Yanfeng Wang, Yu Wang:
Leveraging Diverse Modeling Contexts With Collaborating Learning for Neural Machine Translation. 2100-2111 - Shuo Li, Xiaojun Bi, Tao Liu, Zheng Chen:
Information Dropping Data Augmentation for Machine Translation Quality Estimation. 2112-2124 - Shuoran Jiang, Qingcai Chen, Yang Xiang, Youcheng Pan, Xiangping Wu:
BaSFormer: A Balanced Sparsity Regularized Attention Network for Transformer. 2125-2140 - Morgan Buisson, Brian McFee, Slim Essid, Hélène C. Crayencour:
Self-Supervised Learning of Multi-Level Audio Representations for Music Segmentation. 2141-2152 - Cong Ma, Xu Han, Linghui Wu, Yaping Zhang, Yang Zhao, Yu Zhou, Chengqing Zong:
Modal Contrastive Learning Based End-to-End Text Image Machine Translation. 2153-2165 - Ruiyu Liang, Yue Xie, Jiaming Cheng, Cong Pang, Björn W. Schuller:
A Non-Invasive Speech Quality Evaluation Algorithm for Hearing Aids With Multi-Head Self-Attention and Audiogram-Based Features. 2166-2176 - Ziqiang Zhang, Sanyuan Chen, Long Zhou, Yu Wu, Shuo Ren, Shujie Liu, Zhuoyuan Yao, Xun Gong, Li-Rong Dai, Jinyu Li, Furu Wei:
SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual Data. 2177-2187 - Rui Liu, Berrak Sisman, Guanglai Gao, Haizhou Li:
Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity Rendering. 2188-2201 - Kshitij Mishra, Mauajama Firdaus, Asif Ekbal:
Please Donate to Save a Life: Inducing Politeness to Handle Resistance in Persuasive Dialogue Agents. 2202-2212 - Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, Shogo Seki:
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion With Annealed Langevin Dynamics. 2213-2226 - Florian Schmid, Khaled Koutini, Gerhard Widmer:
Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models. 2227-2241 - Michael Neri, Archontis Politis, Daniel Aleksander Krause, Marco Carli, Tuomas Virtanen:
Speaker Distance Estimation in Enclosures From Single-Channel Audio. 2242-2254 - Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic:
Large-Scale Unsupervised Audio Pre-Training for Video-to-Speech Synthesis. 2255-2268 - Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Ha-Jin Yu:
FA-ExU-Net: The Simultaneous Training of an Embedding Extractor and Enhancement Model for a Speaker Verification System Robust to Short Noisy Utterances. 2269-2282 - Yang Ai, Zhen-Hua Ling:
Low-Latency Neural Speech Phase Prediction Based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks. 2283-2296 - Yanxiong Li, Jialong Li, Yongjie Si, Jiaxin Tan, Qianhua He:
Few-Shot Class-Incremental Audio Classification With Adaptive Mitigation of Forgetting and Overfitting. 2297-2311 - Tianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou Li:
Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification. 2324-2337 - Lei Zhao, Wenbo Zhu, Shengqiang Li, Hong Luo, Xiao-Lei Zhang, Susanto Rahardja:
Multi-Resolution Convolutional Residual Neural Networks for Monaural Speech Dereverberation. 2338-2351 - Christian Geishauser, Carel van Niekerk, Nurul Lubis, Hsien-Chin Lin, Michael Heck, Shutong Feng, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic:
Learning With an Open Horizon in Ever-Changing Dialogue Circumstances. 2352-2366 - Yusuf Eren, Buket Çolak Güvenç, Engin Cemal Mengüç:
Cost-Effective Acoustic Feedback Cancellers for Digital Hearing Aids. 2367-2377 - Jianchen Li, Jiqing Han, Fan Qian, Tieran Zheng, Yongjun He, Guibin Zheng:
Distance Metric-Based Open-Set Domain Adaptation for Speaker Verification. 2378-2390 - Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino:
Masked Modeling Duo: Towards a Universal Audio Pre-Training Framework. 2391-2406 - Guangzhi Sun, Chao Zhang, Philip C. Woodland:
Graph Neural Networks for Contextual ASR With the Tree-Constrained Pointer Generator. 2407-2417 - Bengt J. Borgström, Michael S. Brandstein:
A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement. 2418-2431 - Kun Wei, Bei Li, Hang Lv, Quan Lu, Ning Jiang, Lei Xie:
Conversational Speech Recognition by Learning Audio-Textual Cross-Modal Contextual Representation. 2432-2444 - Shang-Yu Su, Yung-Sung Chung, Yun-Nung Chen:
Joint Dual Learning With Mutual Information Maximization for Natural Language Understanding and Generation in Dialogues. 2445-2452 - Cunhang Fan, Mingming Ding, Jianhua Tao, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Zhao Lv:
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection. 2453-2466 - Hassan Taherian, DeLiang Wang:
Multi-Channel Conversational Speaker Separation via Neural Diarization. 2467-2476 - Sherif Abdulatif, Ruizhe Cao, Bin Yang:
CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement. 2477-2493 - Puhai Yang, Heyan Huang, Shumin Shi, Xian-Ling Mao:
STN4DST: A Scalable Dialogue State Tracking Based on Slot Tagging Navigation. 2494-2507 - Hang Chen, Qing Wang, Jun Du, Bao-Cai Yin, Jia Pan, Chin-Hui Lee:
Optimizing Audio-Visual Speech Enhancement Using Multi-Level Distortion Measures for Audio-Visual Speech Recognition. 2508-2521 - Anderson Queiroz, Rosângela Coelho:
Harmonic Detection From Noisy Speech With Auditory Frame Gain for Intelligibility Enhancement. 2522-2531 - Maodi Hu, Li Qian, Zhijun Chang, Zhixiong Zhang:
KDPG-Enhanced MRC Framework for Scientific Entity Recognition in Survey Papers. 2532-2543 - Leanne Nortje, Dan Oneata, Herman Kamper:
Visually Grounded Few-Shot Word Learning in Low-Resource Settings. 2544-2554 - Purnima Kamath, Chitralekha Gupta, Lonce Wyse, Suranga Nanayakkara:
Example-Based Framework for Perceptually Guided Audio Texture Generation. 2555-2565 - Arka Roy, Udit Satija:
A Novel Multi-Head Self-Organized Operational Neural Network Architecture for Chronic Obstructive Pulmonary Disease Detection Using Lung Sounds. 2566-2575 - Rongzhi Gu, Yi Luo:
ReZero: Region-Customizable Sound Extraction. 2576-2589 - Wenbin Wang, Yang Song, Sanjay K. Jha:
USAT: A Universal Speaker-Adaptive Text-to-Speech Approach. 2590-2604 - Han Han, Vincent Lostanlen, Mathieu Lagrange:
Learning to Solve Inverse Problems for Perceptual Sound Matching. 2605-2615 - Nursadul Mamun, John H. L. Hansen:
Speech Enhancement for Cochlear Implant Recipients Using Deep Complex Convolution Transformer With Frequency Transformation. 2616-2629 - Cheng Peng, Haobo Wang, Jue Wang, Lidan Shou, Ke Chen, Gang Chen, Chang Yao:
Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification. 2630-2640 - Junchuan Zhao, Low Qi Hong Chetwin, Ye Wang:
SinTechSVS: A Singing Technique Controllable Singing Voice Synthesis System. 2641-2653 - Hyung-Seok Oh, Sang-Hoon Lee, Seong-Whan Lee:
DiffProsody: Diffusion-Based Latent Prosody Generation for Expressive Speech Synthesis With Prosody Conditional Adversarial Training. 2654-2666 - Stefano Damiano, Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti:
A Compressive Sensing Approach for the Reconstruction of the Soundfield Produced by Directive Sources in Reverberant Rooms. 2667-2679 - Jiaming Cheng, Ruiyu Liang, Lin Zhou, Li Zhao, Chengwei Huang, Björn W. Schuller:
Residual Fusion Probabilistic Knowledge Distillation for Speech Enhancement. 2680-2691 - Shih-Lun Wu, Chris Donahue, Shinji Watanabe, Nicholas J. Bryan:
Music ControlNet: Multiple Time-Varying Controls for Music Generation. 2692-2703 - Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien:
Contrastive Self-Supervised Speaker Embedding With Sequential Disentanglement. 2704-2715 - Kavya Ranjan Saxena, Vipul Arora:
Interactive Singing Melody Extraction Based on Active Adaptation. 2729-2738 - Shuoran Jiang, Youcheng Pan, Qingcai Chen, Yang Xiang, Xiangping Wu:
Learning to Improve Out-of-Distribution Generalization via Self-Adaptive Language Masking. 2739-2750 - Alexander Shirnin, Nikita Andreev, Sofia Potapova, Ekaterina Artemova:
Analyzing the Robustness of Vision & Language Models. 2751-2763 - Han Ding, Linwei Zhai, Cui Zhao, Fei Wang, Ge Wang, Wei Xi, Zhi Wang, Jizhong Zhao:
Genre Classification Empowered by Knowledge-Embedded Music Representation. 2764-2776 - Lester Phillip Violeta, Ding Ma, Wen-Chin Huang, Tomoki Toda:
Pretraining and Adaptation Techniques for Electrolaryngeal Speech Recognition. 2777-2789
manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.