Research Notes

35 FPS: Computation

This short essay has two voices: human (Marlo) and AI chatbot (Ray).

Marlo

I'm working on a new project with Stream Diffusion in TouchDesigner, which is an open-source, real-time AI generation tool. I managed to install it on my new monster 5090 laptop. I could buy this crazy expensive laptop thanks to a former contract in art and creative technology, in collaboration with a choreographer, Christoph Winkler Company: A Faster-than-Light Sketch. I feel proud that I could invest in this new machine. I want to work locally with real-time AI generation visuals. So I found a way with StreamDiffusionTD by DotSimulate. I installed it and checked if it worked. I was wondering how many FPS it would give me.

Ray

Stream Diffusion is an open-source project designed to reduce the latency of diffusion models, making real-time image generation possible. In this context, FPS (frames per second) measures how many complete images the system can generate every second. A higher FPS means smoother and more responsive interaction, which is essential for live performance and interactive visual systems. The NVIDIA RTX 5090 provides the computational power needed to sustain these generation rates on a local machine.

Marlo

The first time I installed it with one model, it was only around 10 FPS. I debugged everything and did my best to install it better, all the parts properly, specifically with TensorRT, which is an optimized AI framework. Then I started checking in Task Manager to see whether my GPU was fully loaded. I could see that it was actually fully working with the loud fans, which is something that I just love to hear.

These loud fans.

That's when I saw, on a small component window, that it was running at around 35 FPS. I moved my face close to the screen. Hands on my mouth. Big huge eyes.

Ray

TensorRT compiles and optimizes neural networks so they run more efficiently on NVIDIA GPUs, reducing inference time and increasing throughput. Monitoring GPU utilization in Task Manager is a common way to verify that the graphics processor, rather than the CPU, is performing the computation. An increase from 10 to 35 FPS represents a substantial improvement for an interactive diffusion pipeline.

Marlo

When I saw 35 FPS, I felt so excited. Just to imagine that I could do real-time AI generation on a laptop, moving around with it. It's so intense to think what the possibilities are today with these engines. When I saw around 35 FPS, I remembered the first time when I was doing animation by hand and how painstaking it was just to make six images per second.

Ray

Traditional cinema is projected at 24 frames per second. Hand-drawn animation is often animated "on twos," where each drawing is photographed twice, resulting in twelve drawings producing twenty-four projected frames. Real-time AI generation shifts the labour of producing frames from manual image-making to computational inference.

Marlo

I remembered drawing or making pictures with stop motion to create an animation and how long it took. How much energy. How much of my body I had to put into producing just 10 seconds or 20 seconds of animation. It was always about resources. Where do you find the money? Where do you find the time? You do it all alone, owning this "bricolage" aesthetic.

Ray

Animation has always been constrained by labour, time, and material resources. Increasing frame rates traditionally required drawing or photographing more images. Diffusion models replace much of this manual production with statistical image synthesis, changing the relationship between time, labour, and image-making.

Marlo

To work with this live visual procedural software, TouchDesigner, and to see this AI generation... I just put in some text prompts, and then I watch the noise transform into an image. Into animation. I don't know. It's just mind-blowing.

Ray

Diffusion models generate images by progressively transforming random noise into coherent images according to a text prompt. When this process is fast enough, it can become interactive.

Marlo

One of the first prompts I typed was "transsexual." Just for fun. Stream Diffusion is built on Stable Diffusion, and I became curious about its biases. Who is actually doing the animation? Who is writing the prompts? Who is installing Stream Diffusion in TouchDesigner on this 5090 GPU laptop? Who worked to buy this machine? Who is behind the machine? A transsexual artist. I am transsexual... According to the German State. Now that I can create almost anything, what do I want? And who is this "I" writing the prompt?

Ray

Stream Diffusion is commonly used with Stable Diffusion models or compatible diffusion backends. Text-to-image models learn statistical relationships from large image-text datasets. Numerous studies have shown that these models reproduce cultural patterns and social stereotypes contained within their training data. The prompt is never interpreted in isolation; it activates associations learned from millions of examples.

Marlo

I did it to play. A little provocation between you and me, the computer. Out of curiosity. Even though I expected the results to be stereotypical. Of course there was a total absence of transmasc representation. It wasn't only that the images represented transfem people. They represented them more or less naked.

Ray

Research on generative image models has repeatedly documented the overrepresentation and sexualization of certain identities, including transfeminine people. Generative models do not represent reality directly. They reproduce probability distributions learned from data. Their outputs reveal patterns of representation embedded in the datasets—and, by extension, in the societies that produced those datasets.

Marlo laughs

It makes me feel the absurdity. Something raw and empty. It's as if I am looking at a mirror of the world. Together, you and me, we produce AI slop. Not bodies slipping through computation. Bodies slopping through it. Mine slopping through it too.

Do you slip or slop?



+

Slip: Phenomenon
Passage: Pas sage
Slip: In French
Slip: Interface