Feature Selection Methods as Attribute Weighting Schemes for Clustering Corporate Financial Statements
Abstract
Financial clustering plays an important role in uncovering hidden patterns and supporting decision-making in high-dimensional financial statement analysis. However, clustering performance is often constrained by noisy, redundant, and heterogeneous attributes that reduce cluster quality, interpretability, and robustness. Although previous studies in financial analytics have extensively explored clustering algorithms, limited attention has been given to systematically evaluating feature-selection-based attribute weighting strategies for improving clustering effectiveness in complex financial datasets. This study investigates the effectiveness of 6 feature-selection-based weighting strategies for improving financial statement clustering using annual financial data from 604 publicly listed companies on the Indonesia Stock Exchange during 2020-2023. Rather than relying solely on feature filtering, the evaluated methods were utilized to prioritize informative financial attributes and improve clustering structure. Clustering performance was assessed using internal validation metrics, including the Silhouette Coefficient, Dunn Index, Calinski Harabasz Score, and Davies Bouldin Index to evaluate cluster compactness, separation, and overall quality. The results demonstrate that feature-selection-based weighting substantially improves clustering quality compared with the baseline. Fisher Score achieved the strongest overall performance with a Silhouette score of 0.9899, Dunn Index of 1.7238, and Davies Bouldin Index of 0.0034, outperforming the baseline values of 0.8678, 0.3509, and 0.8716, respectively. Autoencoder-based ranking also produced highly competitive results, achieving a Silhouette score of 0.9757 and a Calinski Harabasz Index of 37,105.45. These findings indicate that prioritizing informative financial attributes significantly enhances cluster compactness, separation, and interpretability. The study contributes a comprehensive comparative evaluation of feature-selection-based weighting schemes and provides practical insights for financial segmentation, risk profiling, decision support, and data-driven financial analytics. This study further offers a foundation for developing more robust clustering frameworks for complex financial datasets.
Keywords
Full Text:
PDFReferences
J. Lauer, “Plastic surveillance: Payment cards and the history of transactional data, 1888 to present,” Big Data Soc., vol. 7, no. 1, Jan. 2020, doi: 10.1177/2053951720907632.
A. Markov, Z. Seleznyova, and V. Lapshin, “Credit scoring methods: Latest trends and points to consider,” Nov. 01, 2022, KeAi Communications Co. doi: 10.1016/j.jfds.2022.07.002.
H. Kaur and S. Arora, “Demographic influences on consumer decisions in the banking sector: evidence from India,” Journal of Financial Services Marketing, vol. 24, no. 3–4, pp. 81–93, Dec. 2019, doi: 10.1057/s41264-019-00067-4.
V. Chang, Q. A. Xu, A. Chidozie, and H. Wang, “Predicting Economic Trends and Stock Market Prices with Deep Learning and Advanced Machine Learning Techniques,” Electronics (Switzerland), vol. 13, no. 17, Sep. 2024, doi: 10.3390/electronics13173396.
X. Liu, S. Salem, L. Bian, J. T. Seong, and H. M. Alshanbari, “Application of machine learning algorithms in the domain of financial engineering,” Alexandria Engineering Journal, vol. 95, pp. 94–100, May 2024, doi: 10.1016/j.aej.2024.03.058.
D. G. Cortés, E. Onieva, I. P. López, L. Trinchera, and J. Wu, “Autoencoder-Enhanced Clustering: A Dimensionality Reduction Approach to Financial Time Series,” IEEE Access, vol. 12, pp. 16999–17009, 2024, doi: 10.1109/ACCESS.2024.3359413.
K. Mialkowska, K. Kaczmarczyk, M. Hernes, and M. Dyvak, “Feature Selection for financial data - comparison,” in Procedia Computer Science, Elsevier B.V., 2022, pp. 3041–3050. doi: 10.1016/j.procs.2022.09.362.
P. Chakri, S. Pratap, Lakshay, and S. K. Gouda, “An exploratory data analysis approach for analyzing financial accounting data using machine learning,” Decision Analytics Journal, vol. 7, Jun. 2023, doi: 10.1016/j.dajour.2023.100212.
D. Zhang and L. Zhou, “Discovering Golden Nuggets: Data Mining in Financial Application,” IEEE Transactions on Systems, Man and Cybernetics, Part C (Applications and Reviews), vol. 34, no. 4, pp. 513–522, Nov. 2004, doi: 10.1109/TSMCC.2004.829279.
H. Abdollahi, J. P. Junttila, and H. Lehkonen, “Clustering asset markets based on volatility connectedness to political news,” Journal of International Financial Markets, Institutions and Money, vol. 93, Jun. 2024, doi: 10.1016/j.intfin.2024.102004.
T. Li, G. Kou, Y. Peng, and P. S. Yu, “An Integrated Cluster Detection, Optimization, and Interpretation Approach for Financial Data,” IEEE Trans. Cybern., vol. 52, no. 12, pp. 13848–13861, Dec. 2022, doi: 10.1109/TCYB.2021.3109066.
G. Kou, Y. Peng, and G. Wang, “Evaluation of clustering algorithms for financial risk analysis using MCDM methods,” Inf. Sci. (N. Y)., vol. 275, pp. 1–12, Aug. 2014, doi: 10.1016/j.ins.2014.02.137.
W. Bao, N. Lianju, and K. Yue, “Integration of unsupervised and supervised machine learning algorithms for credit risk assessment,” Expert Syst. Appl., vol. 128, pp. 301–315, Aug. 2019, doi: 10.1016/j.eswa.2019.02.033.
G. Wang, F. Li, P. Zhang, Y. Tian, and Y. Shi, “Data Mining for Customer Segmentation in Personal Financial Market,” 2009, pp. 614–621. doi: 10.1007/978-3-642-02298-2_90.
C. Yang, G. Liu, C. Yan, and C. Jiang, “A clustering-based flexible weighting method in AdaBoost and its application to transaction fraud detection,” Science China Information Sciences, vol. 64, no. 12, Dec. 2021, doi: 10.1007/s11432-019-2739-2.
F. Baser, O. Koc, and A. S. Selcuk-Kestel, “Credit risk evaluation using clustering based fuzzy classification method,” Expert Syst. Appl., vol. 223, p. 119882, Aug. 2023, doi: 10.1016/j.eswa.2023.119882.
S. Dzuba and D. Krylov, “Cluster analysis of financial strategies of companies,” Mathematics, vol. 9, no. 24, Dec. 2021, doi: 10.3390/math9243192.
S. Becirovic, E. Zunic, and D. Donko, “A Case Study of Cluster-based and Histogram-based Multivariate Anomaly Detection Approach in General Ledgers,” in 2020 19th International Symposium INFOTEH-JAHORINA (INFOTEH), IEEE, Mar. 2020, pp. 1–6. doi: 10.1109/INFOTEH48170.2020.9066333.
G. Herman, B. Zhang, Y. Wang, G. Ye, and F. Chen, “Mutual information-based method for selecting informative feature sets,” Pattern Recognit., vol. 46, no. 12, pp. 3315–3327, Dec. 2013, doi: 10.1016/j.patcog.2013.04.021.
T. G. Penkova, “Principal component analysis and cluster analysis for evaluating the natural and anthropogenic territory safety,” Procedia Comput. Sci., vol. 112, pp. 99–108, 2017, doi: 10.1016/j.procs.2017.08.179.
A. Bahl et al., “Recursive feature elimination in random forest classification supports nanomaterial grouping,” NanoImpact, vol. 15, Mar. 2019, doi: 10.1016/j.impact.2019.100179.
C. Barile, C. Casavola, G. Pappalettera, and V. Paramsamy Kannan, “Laplacian score and K-means data clustering for damage characterization of adhesively bonded CFRP composites by means of acoustic emission technique,” Applied Acoustics, vol. 185, p. 108425, Jan. 2022, doi: 10.1016/j.apacoust.2021.108425.
I. Guyon and A. Elisseeff, “An Introduction to Variable and Feature Selection,” 2003.
M. Al Fatih Abil Fida, T. Ahmad, and M. Ntahobari, “Variance Threshold as Early Screening to Boruta Feature Selection for Intrusion Detection System,” in 2021 13th International Conference on Information & Communication Technology and System (ICTS), IEEE, Oct. 2021, pp. 46–50. doi: 10.1109/ICTS52701.2021.9608852.
Y. Siti Ambarwati and S. Uyun, “Feature Selection on Magelang Duck Egg Candling Image Using Variance Threshold Method,” in 2020 3rd International Seminar on Research of Information Technology and Intelligent Systems (ISRITI), IEEE, Dec. 2020, pp. 694–699. doi: 10.1109/ISRITI51436.2020.9315486.
X. He and P. Niyogi, “Locality Preserving Projections,” 2004.
R. Shang, W. Zhang, M. Lu, L. Jiao, and Y. Li, “Feature selection based on non-negative spectral feature learning and adaptive rank constraint,” Knowl. Based. Syst., vol. 236, Jan. 2022, doi: 10.1016/j.knosys.2021.107749.
D. García-García and R. Santos-Rodríguez, “Spectral Clustering and Feature Selection for Microarray Data,” in 2009 International Conference on Machine Learning and Applications, IEEE, Dec. 2009, pp. 425–428. doi: 10.1109/ICMLA.2009.86.
L. Yu and H. Liu, “Efficient Feature Selection via Analysis of Relevance and Redundancy,” 2004.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the Dimensionality of Data with Neural Networks,” Science (1979)., vol. 313, no. 5786, pp. 504–507, Jul. 2006, doi: 10.1126/science.1127647.
R. A. Fisher, “The Use of Multiple Measurements In Taxonomic Problems,” Ann. Eugen., vol. 7, no. 2, pp. 179–188, Sep. 1936, doi: 10.1111/j.1469-1809.1936.tb02137.x.
S.-L. Huang, X. Xu, and L. Zheng, “An Information-Theoretic Approach to Unsupervised Feature Selection for High-Dimensional Data,” IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 1, pp. 157–166, May 2020, doi: 10.1109/JSAIT.2020.2981538.
A. Rizwan, N. Iqbal, A. N. Khan, R. Ahmad, and D. H. Kim, “Toward Effective Pattern Recognition Based on Enhanced Weighted K-Mean Clustering Algorithm for Groundwater Resource Planning in Point Cloud,” IEEE Access, vol. 9, pp. 130154–130169, 2021, doi: 10.1109/ACCESS.2021.3111112.
P. Li, C. Wang, J. Wu, and R. Madlenak, “An E-commerce Customer Segmentation Method based on RFM Weighted K-means,” in 2022 International Conference on Management Engineering, Software Engineering and Service Sciences (ICMSS), IEEE, 2022, pp. 61–68. doi: 10.1109/ICMSS55574.2022.00017.
J.-H. Chen, M.-C. Su, and B. Annuerine Badjie, “Exploring and weighting features for financially distressed construction companies using Swarm Inspired Projection algorithm,” Advanced Engineering Informatics, vol. 30, no. 3, pp. 376–389, Aug. 2016, doi: 10.1016/j.aei.2016.05.003.
A. Kukkar et al., “ProRE: An ACO- based programmer recommendation model to precisely manage software bugs,” Journal of King Saud University - Computer and Information Sciences, vol. 35, no. 1, pp. 483–498, 2023, doi: https://doi.org/10.1016/j.jksuci.2022.12.017.
Madhumita and S. Paul, “A Feature Weighting-Assisted Approach for Cancer Subtypes Identification From Paired Expression Profiles,” IEEE/ACM Trans. Comput. Biol. Bioinform., vol. 19, no. 3, pp. 1403–1414, May 2022, doi: 10.1109/TCBB.2020.3041723.
M. Rafi, H. Khan, H. Nadeem, and H. Shakeel, “Unsupervised Topic Aware Document-Level Semantic Representation for Document Clustering,” in 2021 22nd International Arab Conference on Information Technology (ACIT), IEEE, Dec. 2021, pp. 1–10. doi: 10.1109/ACIT53391.2021.9677217.
A. Aradnia, M. A. Haeri, and M. M. Ebadzadeh, “Adaptive Explicit Kernel Minkowski Weighted K-means,” Inf. Sci. (N. Y)., vol. 584, pp. 503–518, 2022, doi: https://doi.org/10.1016/j.ins.2021.10.048.
K. Balaji, “Machine learning algorithm for feature space clustering of mixed data with missing information based on molecule similarity,” J. Biomed. Inform., vol. 125, p. 103954, 2022, doi: https://doi.org/10.1016/j.jbi.2021.103954.
A. S. Abdulameer, S. Tiun, N. S. Sani, M. Ayob, and A. Y. Taha, “Enhanced clustering models with wiki-based k-nearest neighbors-based representation for web search result clustering,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 3, pp. 840–850, 2022, doi: https://doi.org/10.1016/j.jksuci.2020.02.003.
A. Benkessirat and N. Benblidia, “A novel feature selection approach based on constrained eigenvalues optimization,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 8, Part A, pp. 4836–4846, 2022, doi: https://doi.org/10.1016/j.jksuci.2021.06.017.
N. Donthu, S. Kumar, D. Mukherjee, N. Pandey, and W. M. Lim, “How to conduct a bibliometric analysis: An overview and guidelines,” J. Bus. Res., vol. 133, pp. 285–296, 2021, doi: 10.1016/j.jbusres.2021.04.070.
S. Chowdhury, N. Helian, and R. Cordeiro de Amorim, “Feature weighting in DBSCAN using reverse nearest neighbours,” Pattern Recognit., vol. 137, p. 109314, 2023, doi: https://doi.org/10.1016/j.patcog.2023.109314.
Y. Miao and Y. Xu, “A K-Means-Based Interpolation Algorithm With Lp-Norm and Feature Weighting,” IEEE Access, vol. 12, pp. 96179 – 96192, 2024, doi: 10.1109/ACCESS.2024.3424265.
X. Ai and Q. Ji, “Research on Gas User Clustering Algorithm: Based on PCA and Attribute Weighting,” in Proceedings of the 2022 2nd International Conference on Control and Intelligent Robotics, in ICCIR ’22. New York, NY, USA: Association for Computing Machinery, 2022, pp. 783–788. doi: 10.1145/3548608.3559307.
M. Landauer, F. Skopik, M. Wurzenberger, and A. Rauber, “Dealing with Security Alert Flooding: Using Machine Learning for Domain-independent Alert Aggregation,” ACM Trans. Priv. Secur., vol. 25, no. 3, Apr. 2022, doi: 10.1145/3510581.
S. Solorio-Fernández, J. A. Carrasco-Ochoa, and J. F. Martínez-Trinidad, “A review of unsupervised feature selection methods,” Artif. Intell. Rev., vol. 53, no. 2, pp. 907–948, Feb. 2020, doi: 10.1007/s10462-019-09682-y.
M. Ahmed, R. Seraj, and S. M. S. Islam, “The k-means algorithm: A comprehensive survey and performance evaluation,” Aug. 01, 2020, MDPI AG. doi: 10.3390/electronics9081295.
R. C. de Amorim, “Feature Relevance in Ward’s Hierarchical Clustering Using the L p Norm,” J. Classif., vol. 32, no. 1, pp. 46–62, Apr. 2015, doi: 10.1007/s00357-015-9167-1.
T. Li, A. Rezaeipanah, and E. M. Tag El Din, “An ensemble agglomerative hierarchical clustering algorithm based on clusters clustering technique and the novel similarity measurement,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 6, pp. 3828–3842, Jun. 2022, doi: 10.1016/j.jksuci.2022.04.010.
DOI: https://doi.org/10.47738/jads.v7i3.1410
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)