Current Issue : July-September Volume : 2026 Issue Number : 3 Articles : 5 Articles
Aim: This study compared spectral profile analysis thresholds, speech-in-noise perception, and cerebral asymmetry among Carnatic musicians, Bharatanatyam dancers, and non-trained individuals and examined the influence of training duration on these measures. Method: A total of 105 right-handed adults (18–30 years) with normal hearing were divided into Carnatic musicians (n = 35), Bharatanatyam dancers (n = 35), and non-trained controls (n = 35). Spectral stream segregation was measured using the spectral profile analysis task, and speech-in-noise perception was evaluated using the Kannada QuickSIN under right, left, and binaural conditions. Cerebral asymmetry was derived from the Laterality Index. As data were non-normally distributed, non-parametric tests were used. Results: Significant group differences emerged for spectral profile thresholds, with dancers outperforming musicians and controls. Both trained groups showed superior speech-in-noise performance compared to non-trained individuals across all listening conditions, though no differences were observed between musicians and dancers. Non-trained listeners displayed a clear right-ear advantage, whereas trained groups showed minimal or no hemispheric asymmetry. Training duration negatively correlated with selected spectral profile thresholds in both trained groups and with binaural SNR-50 in dancers, indicating training-related auditory enhancement. Conclusions: Musicians and dancers demonstrate better spectral discrimination, improved speech-in-noise perception, and reduced cerebral asymmetry compared to non-trained peers. These findings underscore training-induced auditory neuroplasticity and suggest that long-term engagement in music or dance promotes efficient auditory processing and greater bilateral hemispheric involvement....
Speaker diarization is a key component for multiple downstream speech technologies, including speech transcription, meeting analytics, and conversational understanding; however, Romanian lacks publicly established diarization resources and benchmarks. This paper evaluates cross-lingual transfer of diarization systems pretrained on predominantly English data, under a strict no-adaptation policy. We compare an end-to-end neural diarization approach (MSDD) and a traditional modular pipeline (segmentation + speaker embeddings + clustering), both used as is with pretrained components. To enable controlled analysis despite the lack of Romanian diarization datasets, we construct a synthetic Romanian conversational benchmark with explicit conditions on speaker count (2–5) and overlap regime (no overlap versus overlap). We report the diarization error rate (DER) and Jaccard error rate (JER) across all conditions, analyze sensitivity to overlap and the number of speakers, and provide an error-component breakdown to identify dominant failure modes. Across all conditions, the end-to-end system outperforms the pipeline (DER 0.140 versus 0.267; JER 0.152 versus 0.320). Performance degrades with overlap and with increasing speaker count in both paradigms, with speaker confusion dominating the additional error under overlap....
This work propose an audio feature pipeline to support machine learning tasks through the extraction of the Mel Frequency Cepstral Coefficients and Mel-spectrogram which is then used as the input of an Convolutional Neural Network which is trained to make the classification tasks. This approach enables de creation of a rich-feature dataset and an end-to-end pipeline that reduces the gap between the audio and Machine Learning ready models with application in sound classification, speech recognition and spatial audio analysis....
Text-to-music generation aims to automatically produce audio content with semantic consistency and coherent musical structure based on natural language descriptions. However, existing methods still face challenges in terms of style diversity, rhythmic consistency, and long-term structural modeling. To address these issues, we propose a novel text-to-music generation model, termed MusicDiffusionNet (MDN), which integrates diffusion models with theWaveNet architecture to jointly model musical semantics and temporal structure in a continuous latent space. By decoupling high-level semantic conditioning from low-level audio generation, MDN enhances its ability to model long-range musical structure while improving semantic alignment between text and generated music with stable generation behavior. Building upon this framework, we further design two complementary mixing strategies to improve generation quality and structural coherence. Adaptive Style Mixing (ASM) performs weighted interpolation among stylistically similar music samples in the style embedding space, incorporating key and harmonic compatibility constraints to expand the style distribution while avoiding dissonance. Multi-scale Temporal Mixing (MTM) adopts beat-aware temporal decomposition, mixing, and reorganization across multiple time scales, thereby enhancing the modeling of both local and global temporal variations while preserving rhythmic periodicity and musical groove. Both strategies are integrated into the diffusion process as conditional augmentation mechanisms, contributing to improved learning stability and representational capacity under limited data conditions. Experimental results on the Audiostock dataset demonstrate that MDN and its mixing strategies achieve consistent improvements across multiple objective metrics, including generation quality, style diversity, and rhythmic coherence, validating the effectiveness of the proposed approach for text-to-music generation....
This research explores the development of a decoder-only speech language model (SLM) for Kazakh, a language currently characterized by limited computational resources. Our approach leverages discrete acoustic units synthesized from self-supervised speech representations. Specifically, we utilize a pretrainedWav2Vec 2.0 model to extract continuous latent features, which are then transformed into discrete semantic tokens via the k-means clustering algorithm. These tokens serve as the foundation for training a generative model designed to predict and maximize the likelihood of speech-unit sequences. To facilitate this study, we curated a specialized Kazakh speech corpus by synthesizing and refining multiple publicly available audio datasets. Given the constrained hardware resources available, we conducted large-scale feature extraction and tokenization to train the unit-based model. We evaluated the system’s efficacy using negative log-likelihood and perplexity metrics on independent test sets. The model captures Kazakh vowel harmony but struggles with long-range agglutinative chains. Key observations include the model’s high sensitivity to data quality, tokenization techniques, and specific training hyperparameters. Although constrained by data volume and training time relative to global benchmarks, the model successfully captures the underlying structural patterns in Kazakh speech. This work establishes a vital empirical baseline and suggests future improvements through refined unit discovery and integrated speech-text modeling....
Loading....