New objective timbre parameters for classification of voice type and fach in professional opera singers

[ad_1]

At the Institute for System Theory and Signal Processing at the University of Stuttgart, a system was developed for classifying professional soprano, tenor, baritone and bass opera voices. To realize this system, various features were developed in a Matlab environment in order to analyze sound samples from lyric and dramatic opera singers within each voice type.

Database

A comprehensive database with a total of 1723 sound examples, all of which can be clearly assigned to a lyric or dramatic category, was compiled and implemented in a Matlab environment. For each sound sample, an analysis interval was defined, containing a single tone with a particular, consistent vowel color. To preserve the original musical context of the investigated tone, every sound sample in the database had a duration of up to 12s. Furthermore, the samples were labeled with codes containing the following information: voice type, voice structure (lyric or dramatic), performer name, pitch, vowel, and title of the musical work. From the 1723 sound samples, 774 were sopranos (352 lyric, 422 dramatic), 389 tenors (199 lyric, 190 dramatic), 343 baritones (157 lyric, 186 dramatic), and 217 basses (60 lyric, 157 dramatic). A more detailed description of this extensive database can be found in a previous study19.

Development and calculation of objective timbre parameters

The parameters for objectifying and specifying the timbre should be able to describe characteristic properties of the spectrum, independent of vowel color. Consequently, a spectral centroid above the vowel formats had to be found. Based on the knowledge of singing voice acoustics, especially the position of vowel formants, as well as the analysis of the extensive sound samples of our database, the following frequency bands (left[{B}_{1},{B}_{2}right]) were determined:

  • soprano 23004500Hz

  • tenor 20003600Hz

  • baritone 20003600Hz

  • bass20003600Hz

For all frequency bands, the center of the sound energy was calculated. To ensure comparability between both voice types and voice structure, the cutoff frequencies were defined to be the same within each gender.

Since different voice types and voice categories produce different sound spectra, this can be used to develop a possible classification. In this study, specific features based on distribution of energy in the SF frequency bands were extracted for quantitative analysis. The energy distribution function (P(f)) is defined as

$$P(f)=frac{sum_{i={B}_{1}}^{f}{S}^{2}(i)}{sum_{i={B}_{1 }}^{{B}_{2}}{S}^{2}(i)}$$

(1)

where (left[{B}_{1},{B}_{2}right]) is the frequency interval corresponding to each voice type.

When the energy distribution function (P(f)) reaches 50% of the whole, the corresponding absolute frequency is defined as Frequency of Half Energy (FHE in Hz), ie (Pleft({f}_{FHE}right)=0.5). The relative position in the frequency interval is defined as Position of Half Energy (PHE in %), which is calculated as

$$PHE=frac{{f}_{FHE} -{B}_{1}}{{B}_{2} – {B}_{1}}*100%$$

(2)

However, PHE or FHE can only represent the midpoint position of the energy distribution and cannot extract further information from the entire SF frequency range. Therefore, this paper defines the energy density function as

$$p(f)=frac{{S}^{2}(f)}{sum_{i={B}_{1}}^{{B}_{2}}{S}^{ 2}(i)}$$

(3)

Since the moments of a function are quantitative measures related to the shape of the function’s graph, the first-order moments of the energy density function are defined as the Spectral Centroid (SC), which is calculated as

$$Centroid(f)=sum_{f={B}_{1}}^{{B}_{2}}f * p(f)$$

(4)

(Centroid(f)) represents the estimator of the energy function in the frequency interval, ie the energetic center of the specified SF frequency range. For the listener, this criterion is essential for the “brightness” of the sound. Similarly, the 2nd, 3rd and 4th order moments are defined as spectral variance, skewness and kurtosis respectively, which also describe SF features.

Statistical analysis

Basic timbre data analysis was performed using descriptive statistics by calculating means and standard deviations (SD) separately for all singers, each voice type (soprano, tenor, baritone, bass) and voice structure (lyric vs. dramatic). To compare the mean values ​​of FHE, PHE and SC depending on voice structure, a two-tailed t-test for independent samples was applied after checking that skewness of the distribution ranged between -1 and 1 in all subgroups. Additionally, we computed reference ranges of FHE, PHE and SC (Mean1.96SD) including voice structure-specific differences in all voice types.

Furthermore, machine learning methods were applied for voice structure classification. Random forest, an excellent method in ensemble learning, was used to significantly improve the accuracy of classification problems. This is achieved by growing an ensemble of base learners and letting them vote for the most popular class22. Base learners in random forest, ie decision trees, randomly select features to increase robustness with respect to noise23. To enhance the accuracy of this model, additional dimensions had to be considered. Besides timbre parameters, the following critical features were extracted for the classification procedure: vibrato parameters (rate of vibrato, VR; extent of vibrato, VE), formant characteristics (strength of formant; start and stop of formant), and perturbation measures (jitter ;shimmer). The ability of the random forest model to analyze the impact of each input feature on the classification was based on information gain during the classification process.

We chose line graphs, correlation heat maps and bar charts as graphical techniques to represent the data. Moreover, the SHapley Additive exPlanations (SHAP), a tool for visualizing machine learning models, was used to explain the output of the training model. All statistical tests were done using SPSS version 26.0.0.1 (IBM, Armonk, NY). The level of significance was set at =0.05. The graphics were created using MATLAB R2017b (The MathWorks, Inc., Natick, MA) and Python version 3.8 (Python Software Foundation, Wilmington, DE). All classification procedures were implemented by Scikit-learn version 1.0.1, the open-source machine learning package in Python.

Sources

1/ https://Google.com/

2/ https://news.google.com/__i/rss/rd/articles/CBMiMmh0dHBzOi8vd3d3Lm5hdHVyZS5jb20vYXJ0aWNsZXMvczQxNTk4LTAyMi0yMjgyMS130gEA?oc=5

The mention sources can contact us to remove/changing this article

[ad_2]

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts