A Hybrid Intelligent Cybersecurity Assessment System For Electronic Document Management Systems
Abstract
Electronic document management systems concentrate confidential and commercially sensitive information behind a single perimeter, making them a high-priority cyberattack target, while existing assessment methods remain largely static, checklist-based, or single-technique, poorly capturing configuration dynamics or prioritizing measures by expected risk reduction. The objective is to develop and validate an intelligent system for the quantitative, explainable cybersecurity assessment of such systems. The novelty and contribution are a hybrid architecture jointly integrating Mamdani fuzzy inference, a stacking ensemble of machine-learning classifiers (random forest, gradient boosting, and a multilayer perceptron), and a Bayesian threat network with an attack graph, unified by a shared domain ontology of document-management assets and threats weighted by confidentiality, integrity, and availability. These heterogeneous estimates are aggregated into a composite cybersecurity assessment index via a fuzzy analytic hierarchy process with adaptive re-weighting from confirmed incidents, while a two-level explainability layer traces each score to its features, fired rules, and probable attack paths. The system was evaluated on 10,239 labelled configuration states from an operational deployment using cross-validation, comparison against seven baseline models, and an ablation study. It achieved the best results among all compared approaches, with an F1-score of 0.946, a Matthews correlation coefficient of 0.927, and an area under the ROC curve of 0.972 – a gain of 2.4 percentage points over the strongest baseline and 10.4 over logistic regression – and the ablation study confirmed that the fuzzy and Bayesian components contribute complementary gains. These findings show that combining data-driven learning with expert-interpretable reasoning yields a more accurate and stable assessment than any single paradigm, and indicate that the framework can support continuous, evidence-based cybersecurity monitoring in document-centric organizations, with future work on adversarial robustness and multi-organization validation.
Keywords
Full Text:
PDFReferences
L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5–32, Oct. 2001. [Online]. Available: https://doi.org/10.1023/A:1010933404324
[A. L. Buczak and E. Guven, "A survey of data mining and machine learning methods for cyber security intrusion detection," IEEE Communications Surveys & Tutorials, vol. 18, no. 2, pp. 1153–1176, Second quarter 2016. [Online]. Available: https://doi.org/10.1109/COMST.2015.2494502
D.-Y. Chang, "Applications of the extent analysis method on fuzzy AHP," European Journal of Operational Research, vol. 95, no. 3, pp. 649–655, Dec. 1996. [Online]. Available: https://doi.org/10.1016/0377-2217(95)00300-2
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: synthetic minority over-sampling technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, Jun. 2002. [Online]. Available: https://doi.org/10.1613/jair.953
T. Chen and C. Guestrin, "XGBoost: a scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, Aug. 13–17, 2016, pp. 785–794. [Online]. Available: https://doi.org/10.1145/2939672.2939785
F. Doshi-Velez and B. Kim, "Towards a rigorous science of interpretable machine learning," arXiv:1702.08608, Feb. 2017. [Online]. Available: https://arxiv.org/abs/1702.08608
T. Fawcett, "An introduction to ROC analysis," Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, Jun. 2006. [Online]. Available: https://doi.org/10.1016/j.patrec.2005.10.010
Forum of Incident Response and Security Teams (FIRST), "Common vulnerability scoring system v4.0: specification document," FIRST.org, Inc., 2023. [Online]. Available: https://www.first.org/cvss/specification-document
M. Frigault, L. Wang, A. Singhal, and S. Jajodia, "Measuring network security using dynamic Bayesian network," in Proc. 4th ACM Workshop on Quality of Protection, Alexandria, VA, USA, Oct. 27, 2008, pp. 23–30. [Online]. Available: https://doi.org/10.1145/1456362.1456368
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA: MIT Press, 2016. [Online]. Available: https://www.deeplearningbook.org/
International Organization for Standardization, "Information security, cybersecurity and privacy protection - guidance on managing information security risks," ISO/IEC 27005:2022, 4th ed., Oct. 2022. [Online]. Available: https://www.iso.org/standard/80585.html
J.-S. R. Jang, "ANFIS: adaptive-network-based fuzzy inference system," IEEE Transactions on Systems, Man, and Cybernetics, vol. 23, no. 3, pp. 665–685, May 1993. [Online]. Available: https://doi.org/10.1109/21.256541
S. Kabir and Y. Papadopoulos, "Applications of Bayesian networks and Petri nets in safety, reliability, and risk assessments: a review," Safety Science, vol. 115, pp. 154–175, Jul. 2019. [Online]. Available: https://doi.org/10.1016/j.ssci.2019.02.009
G. Ke et al., "LightGBM: a highly efficient gradient boosting decision tree," in Advances in Neural Information Processing Systems 30 (NeurIPS 2017), Long Beach, CA, USA, Dec. 4–9, 2017, pp. 3146–3154. [Online]. Available: https://proceedings.neurips.cc/paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf
R. L. Keeney and H. Raiffa, Decisions with Multiple Objectives: Preferences and Value Trade-Offs. Cambridge, England: Cambridge University Press, 1993.
I. Kotenko and A. Chechulin, "A cyber attack modeling and impact assessment framework," in Proc. 5th Int. Conf. Cyber Conflict (CyCon), Tallinn, Estonia, Jun. 4–7, 2013, pp. 119–142.
M. Kuhn and K. Johnson, Applied Predictive Modeling. New York: Springer, 2013. [Online]. Available: https://doi.org/10.1007/978-1-4614-6849-3
S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Advances in Neural Information Processing Systems 30 (NeurIPS 2017), Long Beach, CA, USA, Dec. 4–9, 2017, pp. 4765–4774. [Online]. Available: https://arxiv.org/abs/1705.07874
E. H. Mamdani and S. Assilian, "An experiment in linguistic synthesis with a fuzzy logic controller," International Journal of Man-Machine Studies, vol. 7, no. 1, pp. 1–13, Jan. 1975. [Online]. Available: https://doi.org/10.1016/S0020-7373(75)80002-2
B. W. Matthews, "Comparison of the predicted and observed secondary structure of T4 phage lysozyme," Biochimica et Biophysica Acta (BBA) – Protein Structure, vol. 405, no. 2, pp. 442–451, Oct. 1975. [Online]. Available: https://doi.org/10.1016/0005-2795(75)90109-9
P. Mell, K. Scarfone, and S. Romanosky, "A complete guide to the common vulnerability scoring system version 2.0," FIRST.org, Inc., Jun. 2007. [Online]. Available: https://www.first.org/cvss/v2/guide
National Institute of Standards and Technology, Guide for Conducting Risk Assessments, NIST Special Publication 800-30, Rev. 1, Sep. 2012. [Online]. Available: https://doi.org/10.6028/NIST.SP.800-30r1
J. Pearl, Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. San Francisco, CA: Morgan Kaufmann, 1988.
F. Pedregosa et al., "Scikit-learn: machine learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825–2830, Oct. 2011. [Online]. Available: https://www.jmlr.org/papers/volume12/pedregosa11a/pedregosa11a.pdf
L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, "CatBoost: unbiased boosting with categorical features," in Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Montréal, Canada, Dec. 3–8, 2018, pp. 6639–6649. [Online]. Available: https://proceedings.neurips.cc/paper/2018/hash/14491b756b3a51daac41c24863285549-Abstract.html
R. S. Ross, Managing Information Security Risk: Organization, Mission, and Information System View, NIST Special Publication 800-39, Mar. 2011. [Online]. Available: https://doi.org/10.6028/NIST.SP.800-39
T. L. Saaty, The Analytic Hierarchy Process. New York: McGraw-Hill, 1980.
W. Saffady, Managing Electronic Records: Methods, Best Practices, and Technologies, 5th ed. Lanham, MD: Rowman & Littlefield, 2017.
A. Shameli-Sendi, R. Aghababaei-Barzegar, and M. Cheriet, "Taxonomy of information security risk assessment (ISRA)," Computers & Security, vol. 57, pp. 14–30, Mar. 2016. [Online]. Available: https://doi.org/10.1016/j.cose.2015.11.001
O. Sheyner, J. Haines, S. Jha, R. Lippmann, and J. M. Wing, "Automated generation and analysis of attack graphs," in Proc. IEEE Symposium on Security and Privacy, Oakland, CA, USA, May 12–15, 2002, pp. 273–284. [Online]. Available: https://doi.org/10.1109/SECPRI.2002.1004377
W. Stallings and L. Brown, Computer Security: Principles and Practice, 4th ed. New York: Pearson, 2018.
E. Triantaphyllou, Multi-Criteria Decision Making Methods: A Comparative Study. New York: Springer, 2000. [Online]. Available: https://doi.org/10.1007/978-1-4757-3157-6
D. H. Wolpert, "Stacked generalization," Neural Networks, vol. 5, no. 2, pp. 241–259, Feb. 1992. [Online]. Available: https://doi.org/10.1016/S0893-6080(05)80023-1
L. A. Zadeh, "Fuzzy sets," Information and Control, vol. 8, no. 3, pp. 338–353, Jun. 1965. [Online]. Available: https://doi.org/10.1016/S0019-9958(65)90241-X
DOI: https://doi.org/10.47738/jads.v7i3.1492
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)