The question
Cloning a voice is cheap now. Can the classic hand-crafted speech features (MFCC, LFCC, CQCC) with simple classifiers still tell a real voice from a synthetic one, or does detection need deep learning?
How we tested it
With Mohammadkhair Awwad at Birzeit University, we trained detectors on the ASVspoof 2019 logical-access data and compared 69 systems: hand-crafted feature sets with classical classifiers such as SVMs, and a log-mel LCNN. Each one was tested on attacks it had never seen, on audio passed through codecs it had never heard, and on a different corpus altogether.
What we found
- The log-mel LCNN reached 5.95% equal error rate on thirteen unseen attacks.
- On the cross-corpus test, five simple silence and duration statistics were the best of all 69 systems. Detectors had learned shortcuts in the training data rather than what makes a voice synthetic.
- The real bottleneck is generalisation, not the choice of features.
Equal error rate (EER) is the point where a detector wrongly rejects as many real voices as it wrongly accepts fake ones; lower is better.


