Complete2022–2025
Gmail filter with an experimental spam classifier
A Gmail API filter built with a student team in 2022, extended in 2025 with a TF-IDF and multilayer-perceptron spam classifier trained on the public CEAS-08 email corpus.
- Stack
- Python
- Gmail API
- scikit-learn
- TF-IDF
- MLP
- Docker

Results
| Metric | Value | Note |
|---|---|---|
| Held-out accuracy | 99.55% | CEAS-08, one stratified 80/20 split (31,323 training / 7,831 test messages), no fixed random seed, no baseline comparison. |
| Legitimate mail marked as spam | 15 of 3,462 | |
| Spam missed | 20 of 4,369 |
Case study
The first version, built with a student team in December 2022, is a Dockerized Python tool on the Gmail API that moved mail from blocklisted senders into a quarantine label; it was later reworked into allowlist-based cleanup.
In August 2025 I added a machine-learning classifier: TF-IDF features over the 5,000 most frequent terms, feeding a multilayer perceptron with two hidden layers of 100 and 50 units, trained on the public CEAS-08 email corpus. On one stratified 80/20 split it classified 99.55% of the 7,831 held-out messages correctly, marking 15 of 3,462 legitimate messages as spam and missing 20 of 4,369 spam messages. The classifier reports predicted spam; it does not move mail.
That result comes from a single split of a public corpus, without a fixed random seed or a baseline model, so it describes this experiment rather than performance on a live inbox. The training notebook records the full evaluation.