AI Music Detector
An experimental music classifier that analyzes uploads and microphone recordings using pretrained audio embeddings, with interactive visualizations and transparent model evaluation.

The context
From an interesting idea to an inspectable system.
Music recorded through speakers can sound different from the clean files used to train a model. This project explores the human-versus-AI classification task from audio inspection through a deployed Streamlit app, comparing simple audio features with pretrained embeddings and testing augmentation with noise, echo, and volume changes.
My contribution
What I built
- 01
Integrated a frozen EfficientAT audio encoder with a logistic-regression classifier trained on labeled human and AI music.
- 02
Built an audio pipeline that checks recording quality and selects up to five 20-second sections across a track before combining their embeddings.
- 03
Added microphone recording, file uploads, playback, waveform and sound-map visualizations, and model-performance views in Streamlit.
- 04
Compared clean and simulated-room benchmarks, with training augmentation and artist/reference grouping across dataset splits.
Architecture
How the work moves
Evidence
Documented results
Honest evaluation
Limits and trade-offs
- The evaluation recordings were reused across experiments, so the reported results are a comparison benchmark rather than a fresh blind test.
- Evaluation covers Rock and Electronic music. Performance on real phone recordings, other genres, and unfamiliar generators has not been established.
- The model score is not a calibrated probability of AI authorship, and a prediction is not proof of who made a song.
- On the original evaluation files, the final model detected 63 of 75 AI tracks and mislabeled 8 of 75 human tracks. Music combining human and AI work falls outside the two-label setup.
Inspect the work
Stack and reproduction
- 1Clone the repository and install Python 3.12 and uv.
- 2Run uv sync --locked, then uv run --locked streamlit run app.py. The trained model is included.
- 3Record through the microphone or upload a WAV, MP3, FLAC, or OGG file lasting 10 seconds to 5 minutes, up to 50 MB, then select Analyze recording.