Data Sharing

The MoA Lab Data Repository is a public archive of research datasets released alongside our publications on phishing, web security, and Internet measurement. While the repository is hosted by the MoA Lab at the University of Tennessee, Knoxville, we are also happy to host data for other researchers. The data on this site is restricted to non-commercial, academic use. Contact Doowon Kim doowon@utk.edu with any questions.


Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection Models
USENIX Security 2025

Abstract: We evaluate popular visual similarity-based phishing detection models on 451k real-world phishing websites and find that high accuracy on curated benchmarks does not survive contact with the wild. Attackers evade detection by targeting model pipelines, mimicking benign logos, or simply removing logos altogether, and our adversarial logo manipulations expose further vulnerabilities in several models.

Dataset & Source Code: Code is available at PhishingEval; re-trained models are on OneDrive and must be placed in the directories the code expects. To support reproducibility, we also release eight datasets:

  • apwg451514 — full APWG dataset, July 2021–July 2023 (domains, screenshots, HTML). Access by request form.
  • phishing4190 — sampled subset of the APWG dataset.
  • failed_examples_csv — 6,000 failure cases from the APWG dataset.
  • archive100 — benign dataset, 100 domains.
  • crawl_benign — benign dataset, 110 common brands (used for ablations and manipulations).
  • expand277 / expand277_new / merge277 / merge277_new — reference lists.
  • visible_dataset2 — visible manipulation dataset.
  • perturbated_dataset — perturbation-based manipulation dataset.

Citation

@inproceedings{ji2025evaluating,
  title     = {Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection Models},
  author    = {Ji, Fujiao and Lee, Kiho and Koo, Hyungjoon and You, Wenhao and Choo, Euijin and Kim, Hyoungshick and Kim, Doowon},
  booktitle = {Proceedings of the 34th USENIX Security Symposium},
  year      = {2025}
}
DATASET 2 TITLE
Paper Artifact(s) from INSTITUTION — VENUE YEAR

Abstract: REPLACE WITH ABSTRACT.

DATASET 3 TITLE
Paper Artifact(s) from INSTITUTION — VENUE YEAR

Abstract: REPLACE WITH ABSTRACT.