Audio style conversion using deep learning

Aakash Ezhilan; R. Dheekksha; S. Shridevi

doi:10.6703/ijase.202109_18(5).004

ScienceGate Book Chapters

JOURNAL ARTICLE

Audio style conversion using deep learning

Aakash Ezhilan R. Dheekksha S. Shridevi

Year: 2021 Journal: International Journal of Applied Science and Engineering Vol: 18 (5)Pages: 1-8

DOI: 10.6703/ijase.202109_18(5).004

Get Full-Text PDF Get Analytical Report

Abstract

Style transfer is one of the most popular uses of neural networks. It has been thoroughly researched, such as extracting the style from famous paintings and applying it to other images thus creating synthetic paintings. Generative adversarial networks (GANs) are used to achieve this. This paper explores the many ways in which the same results can be achieved with audio related tasks, for which a plethora of new applications can be found. Analysis of different techniques used to transfer styles of audios, specifically changing the gender of the audio is implemented. The Crowd sourced high-quality UK and Ireland English Dialect speech data set was used. In this paper, the input is the male or female wave form and the opposite gender’s waveform is synthesized by the network, with the content spoken remaining the same. Different architectures are explored, from naive techniques and directly training audio waveforms against convolution neural networks (CNN) to using extensive algorithms researched for image style conversion and generation of spectrograms (using GANs) to be trained on CNNs. This research has a broader scope when used in converting music from one genre to another, identification of synthetic voices, curating voices for AIs based on preference etc.

Keywords:

Computer science Spectrogram Scope (computer science) Convolutional neural network Style (visual arts) Set (abstract data type) Generative grammar Speech recognition Artificial intelligence Artificial neural network Natural language processing Multimedia Art Visual arts

Metrics

Cited By

0.29

FWCI (Field Weighted Citation Impact)

Refs

0.54

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Speech and Audio Processing

Physical Sciences → Computer Science → Signal Processing

Music and Audio Processing

Physical Sciences → Computer Science → Signal Processing

Speech Recognition and Synthesis

Physical Sciences → Computer Science → Artificial Intelligence

Audio style conversion using deep learning

Abstract

Metrics

Citation History

Topics

Related Documents

Audio style conversion using deep learning

Audio style conversion using deep learning

Audio to Text Conversion Using Deep Learning

Font Style Conversion Based on Deep Learning

Conversion of Sign Language to Text and Audio Using Deep Learning Techniques