Date of Award
8-2026
Document Type
Dissertation
Degree Name
Doctor of Philosophy (PhD)
Department
Computer Science
Committee Chair/Advisor
Kai Liu
Committee Member
Long Cheng
Committee Member
Rong Ge
Committee Member
Siyu Huang
Abstract
Machine learning and deep learning have become central to modern computational intelligence, driving breakthroughs across science, engineering, and daily life. In recent years, the volume and complexity of real-world data have increased exponentially due to advances in sensing technologies, large-scale simulations, and digital interactions. As datasets continue to grow in size, dimensionality, and heterogeneity, ensuring the efficiency, scalability, and robustness of learning algorithms has become a critical research challenge. Under such settings, matrix optimization has emerged as a fundamental mathematical framework that provides both theoretical rigor and computational tractability. By leveraging intrinsic matrix structures, such as low-rankness, sparsity, and orthogonality, matrix optimization enables the design of algorithms that can reveal latent structures in high-dimensional data. Moreover, it facilitates the representation of complex data and model parameters on low-dimensional manifolds, enhancing the stability, interpretability, and generalization of machine learning and deep learning models.
In machine learning, many problems can naturally be formulated in matrix form, where both data and model parameters are represented as matrices under various structural constraints. Consequently, related algorithms are expressed as optimization problems over matrix variables that capture relationships within data or model structures. This dissertation investigates several matrix optimization formulations that advance classical learning frameworks through robust and structured formulations. First, incorporating norms such as the $\ell_1$-norm enhances the robustness of Principal Component Analysis (PCA) by reducing the influence of outliers. Second, in multiview learning, multiview nonnegative matrix factorization (NMF) is extended to jointly analyze heterogeneous data sources, producing unified representations that capture both shared and view-specific structures. The Frobenius norm measures the discrepancy between each view-specific matrix and the consensus representation. Third, orthogonality constraints impose stricter independence among rows or columns, preserving maximal information and improving representation quality, leading to more stable and disentangled feature learning. Across these studies, alternating minimization, proximal algorithms, and low-rank approximation techniques are employed to achieve computational efficiency and convergence. These contributions demonstrate how matrix optimization enhances the robustness, interpretability, and scalability of traditional machine learning methods.
Deep learning, as a specialized branch of machine learning, has strengthened its connection with matrix optimization, particularly in the parameterization and training of neural networks. The massive number of parameters in deep models introduces significant challenges in storage, computation, and fine-tuning. Building upon these foundations, this dissertation explores the role of matrix optimization in deep learning, with a focus on Low-Rank Adaptation (LoRA) for pretrained models. LoRA decomposes parameter updates into low-rank matrices, substantially reducing the number of trainable parameters while preserving the expressive capacity of the full model. This research conducts both theoretical and empirical investigations of LoRA and its variants. Several strategies, including Principal Singular Space Adaptation (PiSSA) and random singular value decomposition (SVD), are explored to improve the initialization and fine-tuning efficiency of LoRA. Furthermore, a new three-factor decomposition method, TLoRA, is evaluated to improve flexibility and representational capacity compared to standard LoRA. Extensive experiments across diverse architectures and benchmark datasets demonstrate that low-rank matrix adaptation not only improves training efficiency and memory utilization but also enhances model transferability across tasks. These findings highlight the pivotal role of matrix optimization in developing efficient, adaptable, and scalable deep learning architectures.
Recommended Citation
Cao, Yarui, "Matrix Optimization in Machine Learning and fine-tuning in Deep Learning" (2026). All Dissertations. 4401.
https://open.clemson.edu/all_dissertations/4401