Research project
EntroURL-Bench: trustworthy malicious URL detection.
EntroURL-Bench is a feature-rich benchmark dataset project for
evaluating malicious URL detection methods across modern phishing,
malware, and malicious-web analysis.
VenueISEC 2026 research track
TypeShort paper
StatusPeer-reviewed publication
Authors: Avijit Gayen, Sayan Mondal, Khokan Mondal, and Angshuman
Jana.
Methodology
Build a richer benchmark before trusting the classifier.
Core method
-
Collect and represent URL samples for malicious URL study.
-
Emphasize richer feature representation beyond shallow lexical
signals.
-
Use cleaning, canonicalization, and large-scale web scraping
to capture richer website context.
-
Support reproducible evaluation across phishing, malware, and
malicious web patterns.
-
Frame evaluation around trustworthy comparison, not only raw
model accuracy.
Why it matters
URL classifiers can appear strong when datasets are old, narrow,
or too shallow. EntroURL-Bench focuses on a benchmark structure
that better supports modern malicious-web evaluation.
Architecture diagram
Benchmark construction and evaluation flow
Seed URL datasetCleaning & canonicalizationWeb scrapingStructural / semantic / entropy featuresBenchmark datasetTrustworthy evaluation
Datasets and results
Feature-rich malicious URL benchmark.
Dataset context
The project centers on a benchmark dataset for malicious URL
detection. The ISEC abstract describes enrichment with
structural, semantic, and entropy-based features, including
internal/external hyperlinks, text entropy, and keyword
distributions. A public dataset repository link has not been
added because no permitted public URL is available in the repo.
Results summary
The ISEC 2026 program lists the work as a short paper in the
Research Papers track, scheduled in Research Session 3: Software
Evolution. The page positions EntroURL-Bench as a reproducible
feature-rich benchmark for trustworthy malicious URL detection.
Artifacts
DOI, citation, slides, dataset, and code.