A Text Summarization-Based Similarity Algorithm for Inter-Subchapter Coherence Analysis in Indonesian Academic Documents
Abstract
In academic writing, the logical relationship between subchapters is essential for maintaining document coherence; however, manual evaluation of inter-subchapter consistency is time-consuming and subjective. This study proposes a hybrid text summarization and similarity-based algorithm integrating feature-based extractive summarization with BERT-based neural summarization and semantic similarity measurement using TF-IDF, cosine similarity, and Latent Semantic Analysis (LSA) to analyze inter-subchapter coherence in Indonesian dissertation qualification proposals. The principal novelty lies in operating at the subchapter level rather than the sentence or document level, enabling structural relationship analysis not addressed by existing approaches. Experiments were conducted on 30 Indonesian dissertation qualification proposals split into 21 documents (70%) for training, 3 (10%) for validation, and 6 (20%) for testing, annotated by domain expert evaluators using six quality criteria. Similarity analysis results show that the proposed method identifies strong semantic alignment between logically connected section pairs, with cosine similarity scores reaching 1.00 for the problem formulation - objectives pair and the background -methodology pair on the test set; these perfect scores reflect the structural consistency of academic proposals rather than normalization artifacts. In quality assessment, the model achieves an average exact-match accuracy of 50% against expert evaluations, with per-proposal accuracy ranging from 17% to 83%. The lower overall accuracy is attributed to BERT's tendency to over-predict quality in poorly structured documents, and these findings are reported as exploratory results given the limited test set size (n=6). The proposed framework makes a meaningful contribution toward automated academic writing assessment tools for Indonesian higher education, providing a structured, data-driven approach to evaluating proposal coherence that can serve as a foundation for future large-scale deployment.
Keywords
Full Text:
PDFReferences
G. Salton and M. J. McGill, *Introduction to Modern Information Retrieval*. New York, NY, USA: McGraw-Hill, 1983.
C. D. Manning, P. Raghavan, and H. Schütze, *Introduction to Information Retrieval*. Cambridge, U.K.: Cambridge University Press, 2008, doi: 10.1017/CBO9780511809071.
S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” *Journal of the American Society for Information Science*, vol. 41, no. 6, pp. 391–407, 1990, doi: 10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781, 2013.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in *Proceedings of NAACL-HLT*, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in *Proceedings of EMNLP*, 2019, doi: 10.18653/v1/D19-1410.
S. D. Myla, E. R. Saini, and E. N. Kapoor, “Auto text summarization in natural language processing: Review,” in *Proceedings of IDCIoT*, 2024, doi: 10.1109/IDCIoT59759.2024.10467405.
R. C. Belwal, S. Rai, and A. Gupta, “Extractive text summarization using clustering-based topic modeling,” *Soft Computing*, vol. 27, pp. 1–18, 2023, doi: 10.1007/s00500-022-07346-4.
A. Kaushik, S. H. Attri, and R. S. Jha, “Exploring text summarization techniques: A review of current challenges and future directions,” in *Proceedings of ICDT*, 2024, doi: 10.1109/ICDT61202.2024.10489291.
M. Cao and H. Zhuge, “Grouping sentences as better language unit for extractive text summarization,” *Future Generation Computer Systems*, vol. 110, pp. 600–611, 2020, doi: 10.1016/j.future.2020.04.024.
S. Bano, S. Khalid, N. M. Tairan, and H. A. Khattak, “Summarization of scholarly articles using BERT and BiGRU,” *Journal of King Saud University – Computer and Information Sciences*, vol. 35, no. 8, 2023, doi: 10.1016/j.jksuci.2023.101578.
J. M. Sanchez-Gomez, M. A. Vega-Rodríguez, and C. J. Pérez, “A decomposition-based multi-objective optimization approach for extractive multi-document text summarization,” *Applied Soft Computing*, vol. 91, 2020, doi: 10.1016/j.asoc.2020.106231.
F. Lan, “Hybrid text similarity measurement using TF-IDF and semantic information,” *Advances in Multimedia*, vol. 2022, pp. 1–10, 2022, doi: 10.1155/2022/7923262.
M. Azam, S. Khalid, S. Almutairi, and H. S. M. Bilal, “Current trends and advances in extractive text summarization: A comprehensive review,” *IEEE Access*, vol. 13, pp. 1–20, 2025, doi: 10.1109/ACCESS.2025.3526400.
D. Anggraini *et al*., “Text similarity analysis for academic documents using NLP approach,” *International Journal of Artificial Intelligence*, 2018.
X. Li, Y. Zhang, and H. Chen, “Text similarity measurement based on semantic and statistical features,” *ICIC Express Letters*, vol. 15, no. 5, pp. 123–130, 2021.
J. Wang and P. Li, “A hybrid approach for text classification and similarity analysis,” *ICIC Express Letters*, vol. 16, no. 3, pp. 201–208, 2022.
C. M. Souza, M. R. G. Meireles, and R. Vimieiro, “A multi-view extractive text summarization approach for long scientific articles,” in *Proceedings of the International Joint Conference on Neural Networks (IJCNN)*, 2022, pp. 1–8, doi: 10.1109/IJCNN55064.2022.9892437.
C. Raffel *et al*., “Exploring the limits of transfer learning with a unified text-to-text transformer,” *Journal of Machine Learning Research*, vol. 21, no. 140, pp. 1–67, 2020.
J. Zhang *et al*., “PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization,” in *Proceedings of ICML*, 2020, pp. 11328–11339.
I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long-document transformer,” arXiv preprint arXiv:2004.05150, 2020.
DOI: https://doi.org/10.47738/jads.v7i3.1425
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)