A Multi-Model Framework for Autonomous Schema Discovery and Hybrid Natural Language Generation-Driven Augmented Business Intelligence
Abstract
Traditional Business Intelligence (BI) frameworks rely heavily on data professionals to manually map relational metadata into structured star schemas during the ETL process, creating a significant operational bottleneck in data preparation and downstream interpretation of insights. To address these limitations, this study introduces an augmented analytics framework for autonomous schema discovery and hybrid Natural Language Generation (NLG) driven Augmented BI. In the data representation layer, a weighted hybrid feature fusion mechanism combines structural database metadata with contextual text embeddings produced by a pre-trained Sentence Transformer (all-MiniLM-L6-v2). In the multi-model machine learning layer, a multi-paradigm execution engine combines unsupervised geometric clustering models (K-Means, K-Medoids, DBSCAN) and supervised classifiers (SVM, Random Forest), and performance is evaluated using Leave-One-Out Cross-Validation (LOOCV). The resulting schemas are then dynamically projected into an in-memory OLAP cube. At the downstream insight interpretation layer, a Hybrid NLG engine combines a context-aware, rule-based router with an autoregressive generative decoding mechanism to autonomously produce adaptive, actionable business commentaries triggered by user-driven OLAP exploration states. Experimental results demonstrate that applying linear semantic scaling optimization (α) substantially mitigates statistical semantic blindness and protects the framework from structural schema misclassification. On the E-Commerce dataset, the proposed K-Medoids+Semantic configuration demonstrated topological superiority, achieving a peak Silhouette Score of 0.611 and a compressed Davies-Bouldin Index of 0.514. Meanwhile, on the high-dimensional Superstore dataset, the pipeline maintained high functional flexibility, stabilizing overall classification accuracy up to 94.74%. Furthermore, the downstream Hybrid NLG engine (Template+Generative) demonstrated high factual integrity and linguistic flexibility, achieving a ROUGE-1 score of 0.85, a ROUGE-2 score of 0.82, and a BLEU score of 0.15. This research provides ABI frameworks that enable accelerated executive decision-making through seamless, data-to-insight automation.
Keywords
Full Text:
PDFReferences
B. M. Olukoya, G. O. Ogunleye, P. O. Olabisi, and A. S. Adegoke, “Heterogeneous Ensemble Feature Selection: An Enhancement Approach to Machine Learning for Phishing Detection,” Int. J. Softw. Eng. Comput. Syst., vol. 10, no. 1, pp. 60–74, 2024, doi: 10.15282/ijsecs.10.1.2024.6.0124.
A. I. Gide and A. A. Mu’azu, “A Novel Approach for Addressing IoT Networks Vulnerabilities in Detection and Classification of DoS/DDoS Attacks,” Int. J. Softw. Eng. Comput. Syst., vol. 10, no. 1, pp. 50–59, 2024, doi: 10.15282/ijsecs.10.1.2024.5.0123.
A. Ramadhanu, J. Na’am, G. W. Nurcahyo, and Y. Yuhandri, “Development of affine transformation method in the reconstruction of songket motif,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 12, no. 2, pp. 600–606, 2022, doi: 10.18517/ijaseit.12.2.14069.
L. Aluso and J. O. Enyejo, “Integrating ETL workflows with LLM-augmented data mapping for automated business intelligence systems,” Int. J. Sci. Res. Mod. Technol., vol. 2, no. 11, pp. 76–89, 2023.
M. Bozdemir and M. Bilgin, “Schema Retrieval with Embeddings and Vector Stores Using Retrieval-Augmented Generation and LLM-Based SQL Query Generation,” Appl. Sci., vol. 16, no. 2, p. 586, 2026.
S. Sharma, Z. Li, and Q. Zhang, “Table-to-Text Generation: A Survey on Challenges, Methods, and Directions,” Inf. Fusion, vol. 86, pp. 1–15, 2022, doi: 10.1016/j.inffus.2022.07.003.
A. Tripathi, T. Bagga, and R. K. Aggarwal, “Strategic impact of business intelligence: A review of literature,” Prabandhan Indian J. Manag., vol. 13, no. 3, pp. 35–48, 2020, doi: 10.17010/pijom/2020/v13i3/151175.
N. Prat, “Augmented Analytics,” Bus. Inf. Syst. Eng., vol. 61, no. 3, pp. 375–380, 2019, doi: 10.1007/s12599-019-00589-0.
Z. Ahmad and M. S. M. Sanjudharan, “Augmented Analytics: The Future of Business Intelligence,” Recent Trends Comput. Sci. Softw. Technol., vol. 5, no. 1, pp. 7–13, 2020.
K. Lepenioti, A. Bousdekis, D. Apostolou, and G. Mentzas, “Human-augmented prescriptive analytics with interactive multi-objective reinforcement learning,” IEEE Access, vol. 9, no. ii, pp. 100677–100693, 2021, doi: 10.1109/ACCESS.2021.3096662.
F. Oesterreich, Thuy Duong; Anton, Eduard; and Xu, “Augmenting the Future: An Exploratory Analysis of the Main Resources, Use Cases and Implications of Augmented Analytics,” ECIS 2021 Res. Pap., no. 19, 2021.
M. Mohri, A. Rostamizadeh, and A. Talwalkar, Foundations of machine learning. MIT press, 2018.
M. Mohtasham Moein et al., “Predictive models for concrete properties using machine learning and deep learning approaches: A review,” J. Build. Eng., vol. 63, p. 105444, 2023, doi: https://doi.org/10.1016/j.jobe.2022.105444.
M. Soori, B. Arezoo, and R. Dastres, “Machine learning and artificial intelligence in CNC machine tools, A review,” Sustain. Manuf. Serv. Econ., p. 100009, 2023, doi: https://doi.org/10.1016/j.smse.2023.100009.
V. Srinivasan, S. Santhanam, and S. Shaikh, “Using reinforcement learning with external rewards for open-domain natural language generation,” J. Intell. Inf. Syst., vol. 56, no. 1, pp. 189–206, 2021, doi: 10.1007/s10844-020-00626-5.
T. Liu, W. Wei, and W. Y. Wang, “Table-to-Text Natural Language Generation with Unseen Schemas,” 2019.
C. Koutras et al., “Valentine: Evaluating Matching Techniques for Dataset Discovery,” in 2021 IEEE 37th International Conference on Data Engineering (ICDE), 2021, pp. 468–479. doi: 10.1109/ICDE51399.2021.00047.
N. Reimers and I. Gurevych, “Sentence-{BERT}: Sentence Embeddings using {S}iamese {BERT}-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Nov. 2019, pp. 3982–3992. doi: 10.18653/v1/D19-1410.
J. Z. Huang, M. K. Ng, H. Rong, and Z. Li, “Automated variable weighting in k-means type clustering,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 27, no. 5, pp. 657–668, 2005, doi: 10.1109/TPAMI.2005.95.
R. Cordeiro de Amorim and B. Mirkin, “Minkowski metric, feature weighting and anomalous cluster initializing in K-Means clustering,” Pattern Recognit., vol. 45, no. 3, pp. 1061–1075, 2012, doi: https://doi.org/10.1016/j.patcog.2011.08.012.
DOI: https://doi.org/10.47738/jads.v7i3.1479
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)