User Behavior Analytics for Insider Threat using Transformer Based Approach
DOI:
https://doi.org/10.18486/ijcsnt/15.1.002Keywords:
Anomaly Detection, CERT, Context Generation, Insider Threat Classification, Large Language Model, Llama, SecureBERT, TransformerAbstract
In the past, attacks stemming from the deliberate or inadvertent actions of employees within the organization did not typically worry people. Due to the frequent reports of data breaches involving many organizations and their own employees, businesses are growing increasingly concerned about the need to keep an eye on user’s behavior on the network. In light of this, this study suggests a method for user behavior analytics. It makes use of the CERT (version 4.2) insider threat dataset, a five-log file synthetic dataset developed for insider threat research. Various raw events are provided by these log files. The behavior analysis of every user in this work is carried out based on log sources. The structural behavior provided by CERT, is in structural format, which is converted into the contextual data by fine-tuning the Llama3.1 model. These contextual data along with the synthesized contextual data are used to train the SecureBERT model for classification purpose. The model train accuracy was observed to start from 67% when trained and test accuracy was obtained to be 76% with 86% precision on 2500 data samples. The model was then trained on over 16 thousand data, observing the accuracy started with 92% in first epoch and reached to nearly 99.99%, which shows large language models can perform better and accurate in contextual user behavior analysis in insider threat.
References
Ehsan Aghaei, Ehab Al-Shaer, Waseem Ghassan Tahseen Shadid, and Xi Niu. Automated CVE analysis for threat prioritization and impact prediction. arXiv preprint arXiv:2309.03040, 2023.
Ehsan Aghaei, Xi Niu, Waseem Shadid, and Ehab Al-Shaer. SecureBERT: A domain-specific language model for cybersecurity. In International Conference on Security and Privacy in Communication Systems, pp. 39–56. Springer, 2022. DOI: https://doi.org/10.1007/978-3-031-25538-0_3
Mohammed Nasser Al-Mhiqani, Rabiah Ahmad, Z. Zainal Abidin, Warusia Yassin, Aslinda Hassan, Karrar Hameed Abdulkareem, Nabeel Salih Ali, and Zahri Yunos. A review of insider threat detection: Classification, machine learning techniques, datasets, open challenges, and recommendations. Applied Sciences, 10(15), 2020. DOI: https://doi.org/10.3390/app10155208
Lindauer Brian. Insider Threat Test Dataset. https://doi.org/10.1184/R1/12841247.v1, 2020. Dataset.
Koustav Dutta, Rasmita Lenka, Priya Gupta, Aarti Goel, Janjhyam Venkata, and Janjhyam Ramesh. Swift Diagnose: A high-performance shallow convolutional neural network for rapid and reliable SARS-CoV-2 induced pneumonia detection. EAI Endorsed Transactions on Pervasive Health and Technology, 10, 2024. DOI: https://doi.org/10.4108/eetpht.10.5581
Fredrik Heiding, Bruce Schneier, Arun Vishwanath, Jeremy Bernstein, and Peter S. Park. Devising and detecting phishing: Large language models vs. smaller human models, 2023. DOI: https://doi.org/10.1109/ACCESS.2024.3375882
Fredrik Heiding, Bruce Schneier, Arun Vishwanath, Jeremy Bernstein, and Peter S. Park. Devising and detecting phishing emails using large language models. IEEE Access, 2024. DOI: https://doi.org/10.1109/ACCESS.2024.3375882
Shadi Jaradat, Richi Nayak, Alexander Paz, Huthaifa I. Ashqar, and Mohammad Elhenawy. Multitask learning for crash analysis: A fine-tuned LLM framework using Twitter data. Smart Cities, 7(5):2422–2465, 2024. DOI: https://doi.org/10.3390/smartcities7050095
Dahye Kim, Dongju Park, Honghyun Cho, and Kang. Insider threat detection based on user behavior modeling and anomaly detection algorithms. Applied Sciences, 9:4018, 2019. DOI: https://doi.org/10.3390/app9194018
Xuemei Li and Huirong Fu. SecureBERT and Llama 2 empowered control area network intrusion detection and classification. arXiv preprint arXiv:2311.12074, 2023.
Mario Pérez-Gomariz, Fernando Cerdán-Cartagena, and Jess García. LM-Hunter: An NLP-powered graph method for detecting adversary lateral movements in APT cyber-attacks at scale. Available at SSRN 4807938.
Abir Rahali and Moulay A. Akhloufi. MalBERTv2: Code-aware BERT-based model for malware identification. Big Data and Cognitive Computing, 7(2), 2023. DOI: https://doi.org/10.3390/bdcc7020060
Abir Rahali and Moulay A. Akhloufi. MalBERTv2: Code-aware BERT-based model for malware identification. Big Data and Cognitive Computing, 7(2):60, 2023. DOI: https://doi.org/10.3390/bdcc7020060
Madhu Raut, Sunita Dhavale, Amarjit Singh, and Atul Mehra. Insider threat detection using deep learning: A review. In 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS), pp. 856–863, 2020. DOI: https://doi.org/10.1109/ICISS49785.2020.9315932
Madhu Raut, Sunita Dhavale, Amarjit Singh, and Atul Mehra. Insider threat detection using deep learning: A review. In 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS), pp. 856–863. IEEE, 2020. DOI: https://doi.org/10.1109/ICISS49785.2020.9315932
Balaram Sharma, Prabhat Pokharel, and Basanta Joshi. User behavior analytics for anomaly detection using LSTM autoencoder–insider threat detection. In Proceedings of the 11th International Conference on Advances in Information Technology, pp. 1–9, 2020. DOI: https://doi.org/10.1145/3406601.3406610
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Aarati Kumari Mahato, Roshan Pokhrel, Om Prakash Mahato

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.