Apache Hive and Spark in Cybersecurity Analytics: A Critical Review Using UNSW-NB15 and Comparative Benchmarks
Deborah Ugwuorah1*, Gabriel Eigbe2
Abstract
The exponential surge in cyber threats, fuelled by data volumes projected to reach 175 zettabytes by 2025, exposes traditional defences to breaches, demanding big data analytics for proactive resilience. This critical review evaluates the applications of Apache Hive and Spark in cybersecurity, focusing on UNSW-NB15 and comparatives, such as CICIDS2017 and Bot-IoT, incorporating methodological evidence from surveys, benchmarks, and case studies. Hive’s batch warehousing yields 4.6x auditing speedups and 3-5x I/O reductions via columnar formats, enabling 95-99% anomaly accuracies in federated setups, yet falters in real-time unstructured handling. Spark’s in-memory ensembles achieve 98.8% multi-class accuracy and 98.7% botnet precision, with 2-14x iterative gains over Hadoop, though memory spills and micro-batch latencies limit production scalability. Hybrids with Flink and Kafka deliver 80% time savings and sub-second responses, but simulated imbalances inflate metrics, amplifying biases and FARs. Federated learning counters privacy erosion but risks poisoning with 2.1% FPR overheads. Contributions include framework differentiations, methodological critiques of dataset realism, and optimisations such as adaptive query execution for 150-500% boosts. Gaps in explainability, 5G integration, and cross-tool validations necessitate bias-adjusted, privacy-centric hybrids for equitable defences, urging interdisciplinary benchmarks to bridge lab-production divides amid AI-amplified threats.
Keywords:
Big Data Analytics; Cybersecurity; Apache Hive; Apache Spark; UNSW-NB15 Dataset
![International Journal of Science, Architecture, Technology and Environment [E-ISSN: 3048-8222]](https://i0.wp.com/ijsate.com/wp-content/uploads/2026/05/LOGO-1.png?fit=723%2C680&ssl=1)