Baraa Said العربية Message

Research · With Mohammadkhair Awwad · Birzeit

Can classical speech features still catch AI-generated voices?

Research testing whether classical speech features can still catch AI-generated voices. Across 69 detection systems, generalisation to unseen attacks and codecs, not the choice of features, turned out to be the real bottleneck.

Three spectrograms from the study: genuine speech, a detected fake and a fake the detector missed
Genuine speech, a caught fake and a missed fake
detection systems compared
69
EER of a log-mel LCNN on 13 unseen attacks
5.95%
silence and duration statistics that topped the cross-corpus test
5

The question

Cloning a voice is cheap now. Can the classic hand-crafted speech features (MFCC, LFCC, CQCC) with simple classifiers still tell a real voice from a synthetic one, or does detection need deep learning?

How we tested it

With Mohammadkhair Awwad at Birzeit University, we trained detectors on the ASVspoof 2019 logical-access data and compared 69 systems: hand-crafted feature sets with classical classifiers such as SVMs, and a log-mel LCNN. Each one was tested on attacks it had never seen, on audio passed through codecs it had never heard, and on a different corpus altogether.

What we found

  • The log-mel LCNN reached 5.95% equal error rate on thirteen unseen attacks.
  • On the cross-corpus test, five simple silence and duration statistics were the best of all 69 systems. Detectors had learned shortcuts in the training data rather than what makes a voice synthetic.
  • The real bottleneck is generalisation, not the choice of features.

Equal error rate (EER) is the point where a detector wrongly rejects as many real voices as it wrongly accepts fake ones; lower is better.

Questions about Catching AI-generated voices

Can AI-generated voices be detected?

In this study, detectors did well on attacks similar to their training data: a log-mel LCNN reached 5.95% EER on thirteen unseen attacks. But on a different corpus, five simple silence and duration statistics beat all 69 systems, showing that generalisation is the real problem.

What is equal error rate (EER)?

The operating point where a detector's false-rejection rate equals its false-acceptance rate. A lower EER means a better detector.

Who worked on this research?

Baraa Said and Mohammadkhair Awwad at Birzeit University.

More work

Questions?

Hiring for an AI or full-stack role, or need a system built? One message is enough.