Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Single Image Super-Resolution with Transformer-Based Architectures: A Deep Dive
Published:
Single Image Super-Resolution (SISR) is one of the fundamental problems in low-level computer vision: given a low-resolution (LR) input, reconstruct a high-resolution (HR) output that is visually faithful and perceptually sharp. For years, convolutional neural networks (CNNs) dominated this space. But since the rise of the Vision Transformer (ViT), the field has shifted dramatically — and Transformer-based SR models now hold state-of-the-art results across virtually every benchmark.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
projects
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
Dinh Phu Tran¹, Dao Duy Hung¹, Daeyoung Kim¹
publications
Paper Title Number 1
Published in , 2009
The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.
Trans2Unet: neural fusion for nuclei semantic segmentation
Published in International Conference on Control, Automation and Information Cciences (ICCAIS), 2022
This paper is about transformer-based medical image super-resolution
Recommended citation: Tran, Dinh-Phu, et al. Trans2Unet: neural fusion for nuclei semantic segmentation. 2022 11th international conference on control, automation and information sciences (ICCAIS). IEEE, 2022.
Download Paper
Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution
Published in British Machine Vision Conf. (BMVC) — Glasgow, UK (Rank A), 2024
This paper is about transformer-based image super-resolution
Recommended citation: Tran, Dinh Phu, Dao Duy Hung, and Daeyoung Kim. Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution. 35th The British Machine Vision Conference (BMVC), 2024
Download Paper
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
Published in AAAI Conf. on Artificial Intelligence (AAAI), Social Track — Philadelphia, USA (Rank A*), 2024
This paper is about creating OCR dataset for Vietnamese
Recommended citation: Do, Thao, et al. Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition. Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 39. No. 27. 2025.
Download Paper
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
Published in International Conf. on Computer Vision (ICCV) — Hawaii, USA (Rank A*), 2025
This paper is about Mamba-based for Efficient Video Super-Resolution
Recommended citation: Phu Tran, Dinh, Dao Duy Hung, and Daeyoung Kim. VSRM: A Robust Mamba-Based Framework for Video Super-Resolution. International Conference on Computer Vision (ICCV), 2025
Download Paper
Video Super-Resolution Technique Leveraging Spatial-Temporal Mamba Architecture
Published in Patent — Under review, 2026
A video super-resolution technique leveraging a spatial-temporal Mamba architecture.
Recommended citation: Dinh Phu Tran et al. "Video Super-Resolution Technique Leveraging Spatial-Temporal Mamba Architecture." Patent (under review), 2026.
SAT: Selective Aggregation Transformer for Image Super-Resolution
Published in IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Finding Track — Denver, USA (Rank A*), 2026
A selective aggregation transformer that adaptively fuses multi-scale features for efficient single image super-resolution.
Recommended citation: Dinh Phu Tran, Thao Do, Saad Wazir, Seongah Kim, Seon Kwon Kim, and Daeyoung Kim. "SAT: Selective Aggregation Transformer for Image Super-Resolution." IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
Published in Under review, 2026
Using text as a semantic bridge for parameter-efficient finetuning of audio-visual models.
Recommended citation: Seongah Kim, Dinh Phu Tran, Hyeontaek Hwang, Saad Wazir, Duc Do Minh, and Daeyoung Kim. "TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning." Under review, 2026.
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
Published in Under review, 2026
Adaptive confidence refinement for reliable audio-visual question answering — learning when to answer and when to abstain.
Recommended citation: Dinh Phu Tran, Jihoon Jeong, Saad Wazir, Seongah Kim, Thao Do, Cem Subakan, and Daeyoung Kim. "Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering." Under review, 2026.
Efficient Encoder-Only Context Compression via Marginal Contribution Scoring
Published in International Conf. on Machine Learning (ICML) Workshop — Seoul, South Korea, 2026
Encoder-only context compression using marginal contribution scoring for efficient long-context modeling.
Recommended citation: Thao Do, Dinh Phu Tran, An Vo, Seon Kwon Kim, and Daeyoung Kim. "Efficient Encoder-Only Context Compression via Marginal Contribution Scoring." ICML Workshop, 2026.
FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution
Published in European Conf. on Computer Vision (ECCV) — Malmö, Sweden (Rank A*), 2026
Frequency-guided orthogonal expert learning for robust real-world image super-resolution.
Recommended citation: Minh Son Hoang*, Dinh Phu Tran*, Quyen Nguyen Duc, Dam Hoang Phuong, and Daeyoung Kim. "FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution." European Conference on Computer Vision (ECCV), 2026. (* equal contribution)
MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation
Published in European Conf. on Computer Vision (ECCV) — Malmö, Sweden (Rank A*), 2026
A context-aware gated decoder for robust medical image segmentation.
Recommended citation: Saad Wazir, Patrick Vibild, Dinh Phu Tran, Seongah Kim, and Daeyoung Kim. "MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation." European Conference on Computer Vision (ECCV), 2026.
talks
Super-Resolution and Its Applications
Published:
Invited guest lecture in CS632 (Embedded Operating Systems) at KAIST, covering image and video super-resolution and its real-world applications.
