User Behavior Analytics for Insider Threat using Transformer Based Approach

Authors

  • Aarati Kumari Mahato Institute of Engineering, Tribhuvan University, Nepal
  • Roshan Pokhrel LogPoint Nepal Pvt. Ltd., Nepal
  • Om Prakash Mahato Nepal Telecommunications Authority, Nepal

DOI:

https://doi.org/10.18486/ijcsnt/15.1.002

Keywords:

Anomaly Detection, CERT, Context Generation, Insider Threat Classification, Large Language Model, Llama, SecureBERT, Transformer

Abstract

 In the past, attacks stemming from the deliberate or inadvertent actions of employees within the organization did not typically worry people. Due to the frequent reports of data breaches involving many organizations and their own employees, businesses are growing increasingly concerned about the need to keep an eye on user’s behavior on the network. In light of this, this study suggests a method for user behavior analytics. It makes use of the CERT (version 4.2) insider threat dataset, a five-log file synthetic dataset developed for insider threat research. Various raw events are provided by these log files. The behavior analysis of every user in this work is carried out based on log sources. The structural behavior provided by CERT, is in structural format, which is converted into the contextual data by fine-tuning the Llama3.1 model. These contextual data along with the synthesized contextual data are used to train the SecureBERT model for classification purpose. The model train accuracy was observed to start from 67% when trained and test accuracy was obtained to be 76% with 86% precision on 2500 data samples. The model was then trained on over 16 thousand data, observing the accuracy started with 92% in first epoch and reached to nearly 99.99%, which shows large language models can perform better and accurate in contextual user behavior analysis in insider threat.

References

Ehsan Aghaei, Ehab Al-Shaer, Waseem Ghassan Tahseen Shadid, and Xi Niu. Automated CVE analysis for threat prioritization and impact prediction. arXiv preprint arXiv:2309.03040, 2023.

Ehsan Aghaei, Xi Niu, Waseem Shadid, and Ehab Al-Shaer. SecureBERT: A domain-specific language model for cybersecurity. In International Conference on Security and Privacy in Communication Systems, pp. 39–56. Springer, 2022. DOI: https://doi.org/10.1007/978-3-031-25538-0_3

Mohammed Nasser Al-Mhiqani, Rabiah Ahmad, Z. Zainal Abidin, Warusia Yassin, Aslinda Hassan, Karrar Hameed Abdulkareem, Nabeel Salih Ali, and Zahri Yunos. A review of insider threat detection: Classification, machine learning techniques, datasets, open challenges, and recommendations. Applied Sciences, 10(15), 2020. DOI: https://doi.org/10.3390/app10155208

Lindauer Brian. Insider Threat Test Dataset. https://doi.org/10.1184/R1/12841247.v1, 2020. Dataset.

Koustav Dutta, Rasmita Lenka, Priya Gupta, Aarti Goel, Janjhyam Venkata, and Janjhyam Ramesh. Swift Diagnose: A high-performance shallow convolutional neural network for rapid and reliable SARS-CoV-2 induced pneumonia detection. EAI Endorsed Transactions on Pervasive Health and Technology, 10, 2024. DOI: https://doi.org/10.4108/eetpht.10.5581

Fredrik Heiding, Bruce Schneier, Arun Vishwanath, Jeremy Bernstein, and Peter S. Park. Devising and detecting phishing: Large language models vs. smaller human models, 2023. DOI: https://doi.org/10.1109/ACCESS.2024.3375882

Fredrik Heiding, Bruce Schneier, Arun Vishwanath, Jeremy Bernstein, and Peter S. Park. Devising and detecting phishing emails using large language models. IEEE Access, 2024. DOI: https://doi.org/10.1109/ACCESS.2024.3375882

Shadi Jaradat, Richi Nayak, Alexander Paz, Huthaifa I. Ashqar, and Mohammad Elhenawy. Multitask learning for crash analysis: A fine-tuned LLM framework using Twitter data. Smart Cities, 7(5):2422–2465, 2024. DOI: https://doi.org/10.3390/smartcities7050095

Dahye Kim, Dongju Park, Honghyun Cho, and Kang. Insider threat detection based on user behavior modeling and anomaly detection algorithms. Applied Sciences, 9:4018, 2019. DOI: https://doi.org/10.3390/app9194018

Xuemei Li and Huirong Fu. SecureBERT and Llama 2 empowered control area network intrusion detection and classification. arXiv preprint arXiv:2311.12074, 2023.

Mario Pérez-Gomariz, Fernando Cerdán-Cartagena, and Jess García. LM-Hunter: An NLP-powered graph method for detecting adversary lateral movements in APT cyber-attacks at scale. Available at SSRN 4807938.

Abir Rahali and Moulay A. Akhloufi. MalBERTv2: Code-aware BERT-based model for malware identification. Big Data and Cognitive Computing, 7(2), 2023. DOI: https://doi.org/10.3390/bdcc7020060

Abir Rahali and Moulay A. Akhloufi. MalBERTv2: Code-aware BERT-based model for malware identification. Big Data and Cognitive Computing, 7(2):60, 2023. DOI: https://doi.org/10.3390/bdcc7020060

Madhu Raut, Sunita Dhavale, Amarjit Singh, and Atul Mehra. Insider threat detection using deep learning: A review. In 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS), pp. 856–863, 2020. DOI: https://doi.org/10.1109/ICISS49785.2020.9315932

Madhu Raut, Sunita Dhavale, Amarjit Singh, and Atul Mehra. Insider threat detection using deep learning: A review. In 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS), pp. 856–863. IEEE, 2020. DOI: https://doi.org/10.1109/ICISS49785.2020.9315932

Balaram Sharma, Prabhat Pokharel, and Basanta Joshi. User behavior analytics for anomaly detection using LSTM autoencoder–insider threat detection. In Proceedings of the 11th International Conference on Advances in Information Technology, pp. 1–9, 2020. DOI: https://doi.org/10.1145/3406601.3406610

Downloads

Published

2026-04-30

How to Cite

User Behavior Analytics for Insider Threat using Transformer Based Approach. (2026). International Journal of Communication Systems and Network Technologies, 15(1), 33-43. https://doi.org/10.18486/ijcsnt/15.1.002