publications
Deep Learning, Computer Vision, Diffusion, Image Synthesis, AIGC.
2026
- CSBoRA: A continual learning method for large language models with true orthogonality and reduced forgettingYuyang Liu, Lai-Man Po, Farrell Hung, Zhuohan Wang, Haoxuan Wu, and 4 more authorsPattern Recognition, 2026
Continual learning in large language models (LLMs) faces significant challenges, including catastrophic forgetting and efficient adaptation to new tasks. This paper introduces Continual Standard-Basis Low-Rank Adaptation (CSBoRA), a continual-learning adapter that achieves true orthogonality by construction. CSBoRA assigns each task to a disjoint standard-basis subspace through a fixed one-hot projection-down matrix, yielding regional (sparse) weight updates that isolate task-specific adaptations and better preserve pretrained knowledge. This design eliminates orthogonality regularization, reduces training overhead, and allows higher effective rank under comparable parameter budgets. Our experiments demonstrate that CSBoRA outperforms existing LoRA-based continual learning approaches in training efficiency, robustness, and test performance. CSBoRA achieves true orthogonality, reduces computational overhead, and preserves pre-trained knowledge more effectively. We evaluate CSBoRA on various benchmarks using T5 and LLaMA models. Results show CSBoRA effectively improves performance, mitigates catastrophic forgetting, and reduces training time. This research contributes to efficient and scalable continual learning methods.
- PROrionEdit: Bridging Reference and Source Images for Generalized Cross-Image EditingZeyu Jiang, Lai Man Po, Xuyuan Xu, Yexin Wang, Guoping Gong, and 4 more authorsIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun 2026
Multimodal image synthesis has made significant progress, yet most editing methods still rely on textual instructions, which are less direct than visual guidance. Recently, a new paradigm edits one image using another as reference, enabling more intuitive manipulation through visual exemplars. We formalize this setting as cross-image editing, where a source image is modified under one or more visual references. We propose OrionEdit, a unified framework that regulates editing via symmetric orthogonal subspace disentanglement and reverse-causal attention with information-flow masks enforcing unidirectional latent dependencies. Built on standard diffusion backbones, OrionEdit enables zero-shot multi-reference editing and outperforms open-source baselines, approaching proprietary models in fidelity and disentanglement. The model is available at https://github.com/cityuhkai/OrionEdit.
- arxivThrough the PRISM: Preference Representation in Intermediate States of Video Diffusion ModelsHaoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li, Hongzheng Yang, and 1 more authorJun 2026
Evaluating video generation with clean, pixel-based reward models disconnects evaluation from the noisy diffusion process and incurs massive VAE decoding costs. In this paper, we challenge this paradigm by asking a fundamental question: Can a powerful video generator inherently discriminate preferences directly from noisy latents? To answer this, we introduce \textbfPRISM (\textbfPreference \textbfRepresentation in \textbfIntermediate \textbfStates of Diffusion \textbfModels). PRISM employs a lightweight Query-based Aggregation head with a frozen video diffusion backbone to decode preference signals from noisy latents. Surprisingly, PRISM not only achieves SOTA preference accuracy but also unlocks strong noise-robustness, which enables early-stage Best-of-N sampling. This allows for filtering suboptimal candidates at the very beginning of denoising, drastically reducing computation while boosting video quality. We also reveal a strong positive correlation between a backbone’s generative performance and its inherent evaluative power, enabling self-improving video backbones.
2025
- CVIU
Comprehensive regional guidance for attention map semantics in text-to-image diffusion modelsHaoxuan Wu, Lai-Man Po, Xuyuan Xu, Kun Li, Yuyang Liu, and 1 more authorComputer Vision and Image Understanding, Jun 2025Diffusion models have shown remarkable success in image generation tasks. However, accurately interpreting and translating the semantic meaning of input text into coherent visuals remains a significant challenge. We observe that existing approaches often rely on enhancing attention maps in a pixel-based or patch-based manner, which can lead to issues such as non-contiguous regions, unintended region leakage, eventually causing attention maps with limited semantic richness, degrade output quality. To address these limitations, we propose CoRe Diffusion, a novel method that provides comprehensive regional guidance throughout the generation process. Our approach introduces a region-assignment mechanism coupled with a tailored optimization strategy, enabling attention maps to better capture and express semantic information of concepts. Additionally, we incorporate mask guidance during the denoising steps to mitigate region leakage. Through extensive comparisons with state-of-the-art methods and detailed visual analyses, we demonstrate that our approach achieves superior performance, offering a more faithful image generation framework with semantically accurate procedure. Furthermore, our framework offers flexibility by supporting both automatic region assignment and user-defined spatial inputs as conditional guidance, enhancing its adaptability for diverse applications.
- ESWARMP-adapter: A region-based Multiple Prompt Adapter for multi-concept customization in text-to-image diffusion modelExpert Systems with Applications, Jun 2025
This paper introduces a novel framework for multi-concept customization in text-to-image diffusion models. At its core is a Multiple Prompt Adapter (MP-Adapter) capable of processing multiple image prompts in parallel, extracting features from target concepts and projecting them into the same latent space as the text prompt. This enables simultaneous handling of multiple concepts using just one reference image per concept. To address challenges in fusing multiple concepts with complex interactions, we propose a Region-based Denoising Framework (RDF) that dynamically generates concept-specific regions of interest during inference, allowing spatially decoupled injection of concept features. By integrating the MP-Adapter and RDF, our end-to-end pipeline enables multi-concept customization with intricate occlusions and interactions while preserving concept identities. This approach surpasses current methods by resolving concept conflicts, identity degradation, and occlusion issues, allowing flexible customization without concept-specific retraining. Both qualitative and quantitative evaluations demonstrate that our framework outperforms state-of-the-art approaches in multi-concept customization tasks, while ablation studies validate the effectiveness of each proposed component. This work significantly advances text-to-image generation capabilities for complex, user-defined concept combinations. Code and models will be released at https://github.com/baojudezeze/RMP-Adapter.
- MMS
Multi-SBoRA: regional and non-overlapping weight updates for multi-concept customization of diffusion modelsMultimedia Systems, Oct 2025Customizing diffusion models for multiple concepts remains challenging due to cross-concept interference. This paper introduces Multi-SBoRA, a novel method for customizing diffusion models for multiple concepts. By leveraging orthogonal standard basis vectors, Multi-SBoRA constructs low-rank matrices for LoRA fine-tuning, enabling regional and non-overlapping weight updates that effectively mitigate crosstalk between different concepts. This approach preserves the knowledge embedded in the pre-trained model and reduces interference between customized concepts, thereby ensuring that each concept is learned independently without compromising its integrity. The localized weight updates also reduce computational overhead and enhance model flexibility. Experimental results demonstrate the optimal quantitative performance of Multi-SBoRA, showcasing its efficacy in addressing multi-concept customization while maintaining independence via orthogonality and localized updates, and mitigating crosstalk effects.
2024
- TCSVTSelf-Calibration Flow Guided Denoising Diffusion Model for Human Pose TransferIEEE Transactions on Circuits and Systems for Video Technology, Oct 2024
The human pose transfer task aims to generate synthetic person images that preserve the style of reference images while accurately aligning them with the desired target pose. However, existing methods based on generative adversarial networks (GANs) struggle to produce realistic details and often face spatial misalignment issues. On the other hand, methods relying on denoising diffusion models require a large number of model parameters, resulting in slower convergence rates. To address these challenges, we propose a self-calibration flow-guided module (SCFM) to establish precise spatial correspondence between reference images and target poses. This module facilitates the denoising diffusion model in predicting the noise at each denoising step more effectively. Additionally, we introduce a multi-scale feature fusing module (MSFF) that enhances the denoising U-Net architecture through a cross-attention mechanism, achieving better performance with a reduced parameter count. Our proposed model outperforms state-of-the-art methods on the DeepFashion and Market-1501 datasets in terms of both the quantity and quality of the synthesized images. Our code is publicly available at https://github.com/zylwithxy/SCFM-guided-DDPM.
- ICONIPSBoRA: Low-Rank Adaptation with Regional Weight UpdatesarXiv preprint arXiv:2407.05413, Oct 2024
This paper introduces Standard Basis LoRA (SBoRA), a novel parameter-efficient fine-tuning approach for Large Language Models that builds upon the pioneering works of Low-Rank Adaptation (LoRA) and Orthogonal Adaptation. SBoRA reduces the number of trainable parameters by half or doubles the rank with the similar number of trainable parameters as LoRA, while improving learning performance. By utilizing orthogonal standard basis vectors to initialize one of the low-rank matrices (either A or B), SBoRA facilitates regional weight updates and memory-efficient fine-tuning. This results in two variants, SBoRA-FA and SBoRA-FB, where only one of the matrices is updated, leading to a sparse update matrix ΔW with predominantly zero rows or columns. Consequently, most of the fine-tuned model’s weights (W0+ΔW) remain unchanged from the pre-trained weights, akin to the modular organization of the human brain, which efficiently adapts to new tasks. Our empirical results demonstrate the superiority of SBoRA-FA over LoRA in various fine-tuning tasks, including commonsense reasoning and arithmetic reasoning. Furthermore, we evaluate the effectiveness of QSBoRA on quantized LLaMA models of varying scales, highlighting its potential for efficient adaptation to new tasks. Code is available at this https URL