ECCV 2026

HRDiT Training-Free High-Resolution Image Generation with
Off-the-Shelf Diffusion Transformer Models

Yu Xue1, Haoxuan Qu1, Zhuoling Li1, Hongbin Xu2, Jianxiong Yin3, Simon See3, Hossein Rahmani1, Jun Liu1

1Lancaster University  ·  2South China University of Technology  ·  3NVIDIA AI Tech Centre

Abstract

Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primarily focus on adapting off-the-shelf U-Net-based diffusion models to high resolutions, with limited progress on adapting off-the-shelf Diffusion Transformer (DiT) models despite their strong text-to-image generation capabilities at limited resolutions. In this work, we find two key challenges particularly hindering the application of off-the-shelf DiT models for high-resolution image synthesis in a training-free manner, namely, spatial disorder and long generation time. To address these challenges, we propose a novel method tailored to adapt off-the-shelf DiT models for high-resolution image synthesis. Extensive experiments show the efficacy of our method.

Detail at 4K

HRDiT generates better detail at 4096×4096. Left is the whole image, right is the detail.

A dense hillside city at dusk, thousands of lit apartment windows stacked up the slope, laundry and air conditioners on the balconies A 800 by 800 pixel window of the same image at full resolution: lit windows and balconies
lit windows and balconies
A Moroccan riad courtyard at midday, intricate zellige tilework, carved stucco arches, a still fountain and orange trees A 800 by 800 pixel window of the same image at full resolution: carved stucco and muqarnas
carved stucco and muqarnas
A young woman in an embroidered folk costume standing in a summer field, dense threadwork and beading, soft overcast light A 800 by 800 pixel window of the same image at full resolution: embroidery, stitch by stitch
embroidery, stitch by stitch
A peacock feather in raking light, individual barbs separated, iridescent eyespot shifting from blue to green A 800 by 800 pixel window of the same image at full resolution: separated feather barbs
separated feather barbs

Compared with FLUX.1-dev

FLUX at 1024×1024

Left is a 1024×1024 image from FLUX.1-dev, scaled up; right is HRDiT. Drag to compare.

HRDiT at 4096 by 4096: A dense hillside city at dusk, thousands of lit apartment windows stacked up the slope, laundry and air conditioners on the balconies Stock FLUX.1-dev at 1024 by 1024, enlarged: A dense hillside city at dusk, thousands of lit apartment windows stacked up the slope, laundry and air conditioners on the balconies FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A spiderweb strung with dew at sunrise, each droplet holding an inverted image of the meadow behind Stock FLUX.1-dev at 1024 by 1024, enlarged: A spiderweb strung with dew at sunrise, each droplet holding an inverted image of the meadow behind FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A fine ink drawing of a brass clockwork orrery, dense crosshatching, gears and rings casting shadow on a wooden desk Stock FLUX.1-dev at 1024 by 1024, enlarged: A fine ink drawing of a brass clockwork orrery, dense crosshatching, gears and rings casting shadow on a wooden desk FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A Moroccan riad courtyard at midday, intricate zellige tilework, carved stucco arches, a still fountain and orange trees Stock FLUX.1-dev at 1024 by 1024, enlarged: A Moroccan riad courtyard at midday, intricate zellige tilework, carved stucco arches, a still fountain and orange trees FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A Gothic cathedral rose window seen from inside, stained glass throwing coloured light across worn stone flagstones Stock FLUX.1-dev at 1024 by 1024, enlarged: A Gothic cathedral rose window seen from inside, stained glass throwing coloured light across worn stone flagstones FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: An autumn gorge with a waterfall, wet mossy boulders, birch and maple in full colour, low sunlight through the canopy Stock FLUX.1-dev at 1024 by 1024, enlarged: An autumn gorge with a waterfall, wet mossy boulders, birch and maple in full colour, low sunlight through the canopy FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A coral reef in clear shallow water, branching corals and anemones, a school of small orange fish, sunlight rippling on the sand Stock FLUX.1-dev at 1024 by 1024, enlarged: A coral reef in clear shallow water, branching corals and anemones, a school of small orange fish, sunlight rippling on the sand FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096
HRDiT at 4096 by 4096: A desert canyon at golden hour, layered sandstone walls in orange and violet, a narrow green river far below Stock FLUX.1-dev at 1024 by 1024, enlarged: A desert canyon at golden hour, layered sandstone walls in orange and violet, a narrow green river far below FLUX.1-dev 1024×1024, enlarged HRDiT 4096×4096

FLUX at 4096×4096

Left is FLUX.1-dev run directly at 4096×4096; right is HRDiT, same prompt and seed.

Stock FLUX.1-dev asked for 4096 by 4096 in one pass: a featureless field. HRDiT at 4096 by 4096: A dense hillside city at dusk, thousands of lit apartment windows stacked up the slope, laundry and air conditioners on the balconies
Stock FLUX.1-dev asked for 4096 by 4096 in one pass: a featureless field. HRDiT at 4096 by 4096: A Moroccan riad courtyard at midday, intricate zellige tilework, carved stucco arches, a still fountain and orange trees
Stock FLUX.1-dev asked for 4096 by 4096 in one pass: a featureless field. HRDiT at 4096 by 4096: A young woman in an embroidered folk costume standing in a summer field, dense threadwork and beading, soft overcast light

Method

HRDiT is a training-free framework built from two components, each aimed at one of the two challenges and at the cause we identify behind it. It runs on off-the-shelf weights and needs no fine-tuning.

Off-the-shelf DiT models at 1024x1024 versus direct and I-Max adaptation at 4096x4096.
Off-the-shelf DiT models generate high-quality images quickly at the resolution they were trained on (1024×1024). Pushed to 4096×4096 — either directly or with the DiT-tailored I-Max — they produce spatial disorder (red boxes) and take a very long time (red text).

Spatial Position Alignment (SPA)

Our analysis traces spatial disorder to the limited expressiveness of the positional embedding mechanism once it is stretched to high resolutions: at high token counts it can no longer keep different pairwise positional signals apart after attention. SPA restores that distinction with two complementary operations. Bundle replaces each token index with a bundle index before it enters the positional encoding, which shrinks the set of pairwise signals the mechanism has to separate. Slide then varies where the bundle boundaries fall and averages over the variants, recovering the positional distinctions that bundling alone would erase inside a bundle.

Illustration of the bundle and slide operations on seven token indices.
The bundle and slide operations, with T = 7 tokens and bundle size N = 3. (a) Bundling maps the seven token indices onto three bundle indices. (b) Sliding the bundle boundaries produces N variants of the mapping, so tokens that share a bundle in one variant are separated in another.

Head-adaptive Attention Pruning (HAP)

Profiling shows where the time actually goes: multi-head attention dominates the cost at high resolutions, accounting for over 90% of the generation time for both FLUX and Stable Diffusion 3 at 8K. HAP gives every attention head its own attention scope — the range of image tokens that head is allowed to look at — and prunes the attention computed outside it. The scopes are derived once, offline, from the off-the-shelf model, so generation stays training-free.

Line chart of the share of generation time spent on multi-head attention, rising with resolution for FLUX and Stable Diffusion 3.
Share of the per-image generation time spent on multi-head attention, for off-the-shelf FLUX and Stable Diffusion 3 across resolutions. Attention goes from a minor cost at 1K to the dominant one at 4K and 8K.

BibTeX

@inproceedings{xue2026hrdit,
  title={{HRDiT}: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models},
  author={Xue, Yu and Qu, Haoxuan and Li, Zhuoling and Xu, Hongbin and Yin, Jianxiong and See, Simon and Rahmani, Hossein and Liu, Jun},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}