Following the frames through a Windows VM

By Kitsune

Work notes ·
Linux Virtualization Graphics Debugging
Last updated on

For a long time, I’ve been following the progress of virtio-gpu for Windows Guests. It is getting closer and closer, but… I decided to play with it and see how far I could get.

That was the motivation behind an experiment with Windows graphics acceleration through the Venus/VirtIO-GPU ecosystem, QEMU, and DXVK. My host has an RTX 4090, and I wanted to explore a shared graphics path without handing the entire device over through PCI passthrough.

The interesting part of this story is that getting a graphics path to do something and getting it to feel good are very different milestones.

When accelerated still feels slow

During the investigation, I reported performance around 20 FPS and supplied profiling and trace information. The experience was laggy enough that simply knowing acceleration was involved didn’t answer the useful question: where was the time going?

The experimental stack had several boundaries to investigate. On the guest side, graphics work could involve DXVK and the Windows Venus components. The virtual GPU and QEMU then formed part of the path toward the host graphics stack.

Each boundary introduced another place where work could wait, arrive in bursts, or fail to become a new visible frame. That made it easy to have a plausible theory without having enough evidence to pick one.

A refresh signal isn’t a fresh frame

One of the useful distinctions in the trace discussion was between the cadence of display timing and the cadence of new scanouts.

The analysis of the captured data suggested that timing activity could be close to a normal 60 Hz rhythm while new scanouts arrived much less frequently and unevenly. That gave us a more specific question than “why is the VM slow?”

Was the guest producing new content regularly? Was that content being submitted regularly? Once a scanout reached QEMU, did it wait there, or was the gap earlier in the path?

Those are related questions, but they need different evidence. A smooth timing signal cannot prove smooth delivery of new content. Likewise, gaps in updates from a mostly static desktop are not enough, on their own, to diagnose the performance of continuous animation.

Compare the two timelines

In this example, refresh ticks stay at 60 Hz while the number and spacing of content updates change. Compare regular delivery with uneven delivery, then select the mostly static desktop.

Refresh ticks versus new content

Synthetic one-second timeline. These are illustrative events, not the captured VM trace; the count describes this window, not a benchmark.

60 Hz refresh ticksNew content arrives0 ms1,000 ms

60 refresh ticks, 20 content updates. Largest gap between updates: 85 ms.

The refresh cadence stays the same while the delivery of new content changes.

These are synthetic timelines, not a replay of the captured trace. The static case is there because few updates can be normal when little is changing. I would want continuous animation before interpreting sparse updates as a performance failure.

Narrowing the question

The working interpretation moved toward irregular delivery before QEMU’s handling of the scanout. I treat that as a direction for investigation, rather than a proven root cause.

A proposed next test was a continuously animated D3D11 workload. That would help separate a desktop that simply had little new content from a graphics path struggling to produce or deliver it.

I still don’t have a polished, reliable Windows GPU-sharing setup to announce from this experiment, but what made the work worthwhile was the shift in how I framed the problem: “GPU acceleration works” was too broad to describe the experience. Following the timing of actual content gave me a way to ask smaller questions and decide what the next trace needed to show.

Sometimes the progress in a debugging session is a fix. Here, it was getting closer to measuring the thing I actually cared about: a new frame arriving when I expected it to.

© 2026 Kitsune