Research project

EntroURL-Bench: trustworthy malicious URL detection.

EntroURL-Bench is a feature-rich benchmark dataset project for evaluating malicious URL detection methods across modern phishing, malware, and malicious-web analysis.

VenueISEC 2026 research track
TypeShort paper
StatusPeer-reviewed publication

Authors: Avijit Gayen, Sayan Mondal, Khokan Mondal, and Angshuman Jana.

Methodology

Build a richer benchmark before trusting the classifier.

Core method
  • Collect and represent URL samples for malicious URL study.
  • Emphasize richer feature representation beyond shallow lexical signals.
  • Use cleaning, canonicalization, and large-scale web scraping to capture richer website context.
  • Support reproducible evaluation across phishing, malware, and malicious web patterns.
  • Frame evaluation around trustworthy comparison, not only raw model accuracy.
Why it matters

URL classifiers can appear strong when datasets are old, narrow, or too shallow. EntroURL-Bench focuses on a benchmark structure that better supports modern malicious-web evaluation.

Architecture diagram

Benchmark construction and evaluation flow

Seed URL datasetCleaning & canonicalizationWeb scrapingStructural / semantic / entropy featuresBenchmark datasetTrustworthy evaluation

Datasets and results

Feature-rich malicious URL benchmark.

Dataset context

The project centers on a benchmark dataset for malicious URL detection. The ISEC abstract describes enrichment with structural, semantic, and entropy-based features, including internal/external hyperlinks, text entropy, and keyword distributions. A public dataset repository link has not been added because no permitted public URL is available in the repo.

Results summary

The ISEC 2026 program lists the work as a short paper in the Research Papers track, scheduled in Research Session 3: Software Evolution. The page positions EntroURL-Bench as a reproducible feature-rich benchmark for trustworthy malicious URL detection.