Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

Single Image Super-Resolution with Transformer-Based Architectures: A Deep Dive

15 minute read

Published:

Single Image Super-Resolution (SISR) is one of the fundamental problems in low-level computer vision: given a low-resolution (LR) input, reconstruct a high-resolution (HR) output that is visually faithful and perceptually sharp. For years, convolutional neural networks (CNNs) dominated this space. But since the rise of the Vision Transformer (ViT), the field has shifted dramatically — and Transformer-based SR models now hold state-of-the-art results across virtually every benchmark.

portfolio

projects

publications

Paper Title Number 1

Published in , 2009

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

Trans2Unet: neural fusion for nuclei semantic segmentation

Published in International Conference on Control, Automation and Information Cciences (ICCAIS), 2022

This paper is about transformer-based medical image super-resolution

Recommended citation: Tran, Dinh-Phu, et al. Trans2Unet: neural fusion for nuclei semantic segmentation. 2022 11th international conference on control, automation and information sciences (ICCAIS). IEEE, 2022.
Download Paper

Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition

Published in AAAI Conf. on Artificial Intelligence (AAAI), Social Track — Philadelphia, USA (Rank A*), 2024

This paper is about creating OCR dataset for Vietnamese

Recommended citation: Do, Thao, et al. Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition. Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 39. No. 27. 2025.
Download Paper

VSRM: A Robust Mamba-Based Framework for Video Super-Resolution

Published in International Conf. on Computer Vision (ICCV) — Hawaii, USA (Rank A*), 2025

This paper is about Mamba-based for Efficient Video Super-Resolution

Recommended citation: Phu Tran, Dinh, Dao Duy Hung, and Daeyoung Kim. VSRM: A Robust Mamba-Based Framework for Video Super-Resolution. International Conference on Computer Vision (ICCV), 2025
Download Paper

SAT: Selective Aggregation Transformer for Image Super-Resolution

Published in IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Finding Track — Denver, USA (Rank A*), 2026

A selective aggregation transformer that adaptively fuses multi-scale features for efficient single image super-resolution.

Recommended citation: Dinh Phu Tran, Thao Do, Saad Wazir, Seongah Kim, Seon Kwon Kim, and Daeyoung Kim. "SAT: Selective Aggregation Transformer for Image Super-Resolution." IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

Published in Under review, 2026

Using text as a semantic bridge for parameter-efficient finetuning of audio-visual models.

Recommended citation: Seongah Kim, Dinh Phu Tran, Hyeontaek Hwang, Saad Wazir, Duc Do Minh, and Daeyoung Kim. "TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning." Under review, 2026.

Efficient Encoder-Only Context Compression via Marginal Contribution Scoring

Published in International Conf. on Machine Learning (ICML) Workshop — Seoul, South Korea, 2026

Encoder-only context compression using marginal contribution scoring for efficient long-context modeling.

Recommended citation: Thao Do, Dinh Phu Tran, An Vo, Seon Kwon Kim, and Daeyoung Kim. "Efficient Encoder-Only Context Compression via Marginal Contribution Scoring." ICML Workshop, 2026.

FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution

Published in European Conf. on Computer Vision (ECCV) — Malmö, Sweden (Rank A*), 2026

Frequency-guided orthogonal expert learning for robust real-world image super-resolution.

Recommended citation: Minh Son Hoang*, Dinh Phu Tran*, Quyen Nguyen Duc, Dam Hoang Phuong, and Daeyoung Kim. "FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution." European Conference on Computer Vision (ECCV), 2026. (* equal contribution)

MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation

Published in European Conf. on Computer Vision (ECCV) — Malmö, Sweden (Rank A*), 2026

A context-aware gated decoder for robust medical image segmentation.

Recommended citation: Saad Wazir, Patrick Vibild, Dinh Phu Tran, Seongah Kim, and Daeyoung Kim. "MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation." European Conference on Computer Vision (ECCV), 2026.

talks

Super-Resolution and Its Applications

Published:

Invited guest lecture in CS632 (Embedded Operating Systems) at KAIST, covering image and video super-resolution and its real-world applications.

teaching