← All Articles

Brief Practical Advice for Short-Form Brainrot

3,665 words

I’m often asked why brainrot works and how to create it. It’s a difficult subject to approach, because there is much to be said and most of it lands as abstract nonsense. Direct practical advice quickly becomes outdated and often misses the core point of why it works. This is an attempt to compress the core ideas into something that can get you started.

Meaning Exists Between Videos

Digital culture moves fast. More content exists than any single person can ever hope to consume, spread across different platforms in multiple languages. Although there are shared threads that unite global culture, the fragmented nature of content distribution makes it so ideas spread and develop into niches that are opaque and hard to parse. Any person, symbol or idea can go from familiar to incomprehensible in a matter of days.

Take, for example, Norwegian soccer player Erling Haaland. During the 2026 FIFA World Cup, he exploded in popularity. His image became the common symbol for an explosion of content. There were the obvious uses of his image, like highlights from his matches and interviews. But then, his image started to differentiate itself outward to narrower and narrower contexts. Jokes that compared his appearance to mops and a green onion’s roots enjoyed global success, as they required no additional context. Caricatures exaggerating his physical traits evolved until he sat somewhere between a monster and a mythical creature, and what that looked like depended on the country of origin. His likeness was adapted to viral memes from China’s own social networks. He started to appear in fake posters for popular franchises, like the latest HBO Game of Thrones adaptation and Love and Deepspace, the most popular Chinese gacha action / dating game.

Side-by-side comparison of Erling Haaland and a green onion's rootsFake Game of Thrones poster featuring Haaland, captioned "Haaland is coming"Fake Love and Deepspace promotional card featuring Haaland
Haaland leek, in Game of Thrones and Love and Deepspace.

At some point, Haaland stopped being a person and became a bridge between different symbols. Even the most plugged-in and in-the-know online users can only hope to situate themselves by the relationship between pieces of content and not all the content itself. I might not know anything about soccer or any of the niche references Haaland was injected into, but I recognize the other elements that were remixed. That allows me to make sense of the immense deluge of new media that enters circulation every day.

Brainrot, slop, memes, or whatever else you want to call it, operates under this logic. It thrives on the correct combination of symbols. What might look like an absurd mixture from the outside often makes perfect sense to a group of people who are able to decode its parts from the whole. This style is a payload of symbolic combinations that people with shared context can quickly decode socially, and people without the proper context can use to bridge to the in-group. Brainrot is a very effective form of social meaning compression.

I avoid using “meme” as a term because it has become both broad and historically constrained. Technically, a meme can be any single piece of media. An image macro, a joke song or a video remix are all “memes,” but Google search interest in the term has declined since 2020, and I mostly associate it with trends from the 2010s. Brainrot started being used disparagingly to refer to low-quality or low-value content, particularly on short-form platforms. With the advent of generative AI media, “AI slop” or just “slop” started to be used in parallel and interchangeably with “brainrot.”

Google Trends chart for the search term memes, worldwide, 2004 to present
Worldwide Google search interest in “memes”, 2004–present.
Google Trends chart for the search term brainrot, worldwide, 2004 to present
Worldwide Google search interest in “brainrot”, 2004–present. Each chart is normalized separately.

So why does this matter? Because the core point to understand about this style is that you have to look between symbols, and not at them. To understand most viral trends, you need to understand that they are proxies for different referential points.

Practical Advice

1. Building Blocks

There are three main axes that can predict the success of brainrot: novelty, sensation and recognition.

Three overlapping circles: Recognition, Sensation, and Novelty. Recognition with Sensation produces nostalgia edits; Recognition with Novelty produces character variations; Novelty with Sensation produces edits that go hard.

Novelty is a matter of degree. How different is this piece of content from everything else I’ve seen? If it’s too different, I have to make an effort to understand it and that makes it difficult for me to engage. If it’s too similar, I already understand it so I don’t need to take the time to engage.

Sensation refers to the audiovisual aspects of the piece. How does it make me feel? Is the music good, and do the visuals match the beat in a way that’s satisfying to watch? Or is the song so bad I become irritated and I can’t believe what my ears are hearing? Music and video are by themselves very powerful, but the correct combination of both can by itself elicit emotion from the viewer, enough to move them in different ways. That is the entire culture of “edits” in short-form video, and especially TikTok.

Recognition is the guarantee that the viewer has enough context to situate the content. Do I recognize a character, a style or the references this content points to? If a viewer has nothing to grasp at, they’ll look away. Creators overestimate how much people care about original creations and originality as a whole.

Tung Tung Tung Sahur beside a similarly shaped rectangular character with a human face
Recognition drives humor.

Pairing two categories from the triad at a time is effective by itself.

Recognition and Sensation: nostalgia edits. Hey, remember this thing? Remember how it made you feel?

This recognition is not constrained to nostalgia, though it is one of the most potent forms of this combo. Millennials remembering bits and pieces of the 80s and 90s ignited trends like vaporwave and synthwave. Gen Z has a similar relationship to Frutiger Aero as a pastiche of styles that mixes different design trends from the late 90s and early 00s, expressed through consumer electronics, videogames and music. The recognition can transport you to a different time, and a different sensation.

A nostalgia edit combining a galaxy hoodie, sunglasses, palm trees and a Snapchat dog filter
Nostalgia drives Recognition and Sensation.

Novelty and Recognition: character variations. We took this popular character and made them in a different style or from a different franchise! Homer Simpson dressed as Peter Griffin from Family Guy. This is the bulk of early internet humor.

Homer Simpson and Peter Griffin exchanging character designs
Character swaps are a cheap form of novelty that maintains recognition.

Novelty and Sensation: edits that go hard. Usually this comes in the form of movie or anime footage set to music genres that never existed in the original. A single strange Spider-Man edit cut to distorted hyperpop spawned an entire “polyester” style.

Opening motion edit before the scene transition
Edits are kinetic toys in the way they match motion to sound feedback, allowing for ample exploration of novelty and sensation. Edit by isshin.
The first two seconds of the Polyester Spider-Man edit
Polyester Spider-Man

Each successive iteration on a hit has diminishing returns. The way to maximize your chances of a breakout success is to look at what connections already exist but haven’t been put together into a single frame yet. That can be the jokes and images being posted in the comment section of a popular video, or maybe two parallel viral trends that haven’t been connected yet.

One of my most popular videos took an existing remix of the song “Hound Dog.” It uses a remix created by user @yup.bro.beatz, which he called an “Elvis Presley Minimal Tech Edit.” The remix itself is made to be grating on purpose, and the large negative reaction exploded into virality when another user, @iheartitsreal, posted a video with the caption “the song of the summer,” with those famous and somewhat grotesque Elvis Presley caricatures overlaying the video.

Original @yup.bro.beatz post from TikTok
Original @yup.bro.beatz post from TikTok
Original @iheartitsreal post from TikTok
Original @iheartitsreal post from TikTok

I just mixed these existing elements together with an added layer of novelty. The caricatures themselves were now singing the song, except they were glossy and rubbery-looking 3D models, with backgrounds mimicking festival DJ visuals. Viewers could recognize the current trending meme; they felt the sensation of how grating the song was and a sense of unease at the visuals, while looking at something unlike anything else available for the meme. Why is this Elvis caricature “oiled up”? Why is the song so bad? Is this a joke?

Viewers found their own meaning and remixed on top of that. Some linked the visuals to the type of visuals you would find in a bowling alley or on the slots. All three elements created a package that could serve as both the basis for further remixing and something you could easily point to that represented the entire joke in a compressed form.

Hound Dog Test 1 on TikTok
Hound Dog Test 1, over 11 million views accumulated in 5 months.

2. Timing

Because things move so fast, the window to make a joke land is short. You need a critical mass of people to see whatever you made within that window, and that requires enough shared context. The successful combination of novelty, recognition and sensation will fall flat if you are too late to the joke, and no one will get it if you’re so early that the conversation hasn’t moved there yet.

This is the hardest thing to teach. The best piece of advice I can give is: if you have to stop and think, it’s already too late. The jokes should land in your head as soon as your unconscious mind recognizes the connection. Successful creators seldom plan the jokes themselves and spend their time making sure they are prepared to act once the idea arrives. That means they hone their practical abilities in editing, filming, scripting, etc.

“I’ll do it tomorrow”—the joke is now dead and you missed your shot.

3. Tools

I’m often asked what tools I use, or what workflows are best. I believe this question is misguided. Tools are always changing and evolving.

The visual language of brainrot evolves from the tools that create it, not from design decisions. These videos often look low quality and disjointed because they were made to be as cheap and easy to create as possible. That means they are made with whatever tools you have access to. Unedited AI outputs, free editors with watermarks, stolen footage only slightly remixed. This is a core component of what makes the style work in the first place. The original Tung Tung Tung Sahur used the default “Adam” voice from ElevenLabs and an unedited output from OpenAI’s DALL-E image models.

If you make brainrot look “too good,” it loses its plausibility. If it looks like a brand’s marketing department made something with a budget, it loses its luster. It makes things look thoughtfully engineered instead of scraped together from whatever you could connect from free alternatives. “So bad it’s good” is a very useful heuristic for brainrot.

Outdated tools can usually give you a nicer uncanny feeling that helps sell the concept of brainrot. You don’t need to worry about using the latest and greatest.

Copyright infringement and reusing popular media are one of the core reasons why brainrot works. It hijacks familiarity and makes engagement with your content a little bit easier. SpongeBob talking about geopolitics? That’s both funny and recognizable.

For that reason, a lot of footage gets sourced directly from big platforms like YouTube or the Internet Archive. The Internet Archive is a fantastic repository of old documentaries, movies and TV shows, including plenty of material in the public domain. Things like SpongeBob and Family Guy tend to be favored because there are so many clips already cut and pasted around that you can just get a Best Family Guy compilation from YouTube and use that as your canvas.

4. Retention Mechanisms

Retention mechanisms are strategies that give viewers a reason to stick around and watch a video. This is the broadest area of study when it comes to what makes people keep watching, and the hardest to distill into a few bullet points.

What’s important to understand is that consumption of short-form content exists in two distinct modes. There is scrolling through the feed, where everything melds into a single flow of action. The videos matter less as individual pieces than as a sequence that never lets you know what comes next. You might see a cat, then some food, then a beautiful person before arriving back at a cat. The feed is essentially infinite and is a retention mechanism for the app itself. You might send interesting videos to your friends, or you might just keep scrolling until you find something that piques your interest.

The second mode is stopping to watch an individual video. Here, you are stepping into a particular subject, creator or meme and choosing to focus your attention rather than continue scrolling. This constitutes a pattern interrupt and is “expensive” relative to the almost meditative state of just scrolling. This is why creators obsess over “hooks”: the short window from the first frame to the first few seconds that convinces a user to stick around and engage with the content. Once a user starts watching, it falls on the video to offer enough reason for them to keep engaging.

Hooks are the subject of endless discussion and exploration in their own right. For our purposes, I want to focus on a few popular strategies to keep people watching after they stop to look. There are many more, but these are worth considering as a starting point.

Stimulus Layering

You are giving your audience multiple channels to engage with, so that if one loses their attention, another is there to catch it. The infamous Subway Surfers gameplay footage or clips from copyrighted shows like Family Guy fit here. Neither strategy is as common as it once was, but the principle remains: a parallel content stream that provides motion, builds expectation or illustrates the narrative.

Captions give you the option to read rather than listen to what someone has to say, and sound effects can cue motion or a change of subject.

A podcast clip above Subway Surfers gameplay
Clips from podcasts can be supplemented with a variety of overlays and parallel footage to keep you watching while you listen.

Pattern Interrupts

Cuts, zooms, camera pans and scene changes shift your point of focus faster than you can habituate. A good video dazzles you into not knowing what comes next. If I see someone talking to a camera with subtitles, I know what the shape of that video will be, so retention falls on the quality of the message or narrative being transmitted. If instead the video cuts to another scene or camera angle, that alone adds a small element of novelty that can keep me from guessing what’s next.

A densely layered editing timeline labeled Would You Risk Drowning for $500,000? by MrBeast
An editing timeline for a MrBeast video. The channel uses extensive micro-editing to keep viewers engaged throughout the video, with frequent camera changes, sounds and motion.

Prediction Errors and the Bizarre

Similar to pattern interrupts, except you never quite grasp what’s happening. “Never let them know your next move” is a popular style of video that lends itself well to the non-sequitur, absurdist nature of brainrot and short-form video as a whole. These videos mix different genres as meta-commentary and follow a single narrative thread that keeps you watching to answer “what is this?” but never resolves into a stable or coherent story. Scenes are connected by action and continuity but not by meaning. A character commits an act of violence, only for it to go ignored in favor of another story beat. These videos achieve something between a dream and a horror story, leaving viewers uncomfortable, laughing or horrified.

Tung Tung Tung Sahur building a house inside a head
Brainrot characters building houses inside each other is another popular, quasi-cartoon style, combining time-lapse construction videos, character drama and bizarre characters.
Patrick Star in a suit preparing food in an AI-generated TikTok video
Videos of popular IP characters reimagined as gangsters or in different professions attract attention. Patrick Star chopping people up to start new fast-food restaurants was a popular genre for a while.

In-group Lexicon as Bait

In these videos, the symbols and vocabulary become a checkpoint for engagement. If you get the reference, you feel addressed and can share it with your friends. If you don’t, you engage to try to understand. The videos become temporary in-group filters that lose effectiveness as more and more people get in on the joke.

TikTok caption saying you need 700-plus hours of screen time to understand this referenceTikTok caption saying you need 40 hours of screen time to understand this
“How it feels to get a message” sits at the heart of why brainrot works: a payload you and your peers can deconstruct through the social activity of decoding shared meaning.

Make a Lot of Content

The way that can be taught is not the way. Exhaustive analysis of what makes short-form video work (and more specifically brainrot) can help, but all the components I discussed are best integrated through practice.

The entire point of the style is that it’s accessible and cheap to produce. Guides and write-ups can satisfy curiosity or help direct your efforts, but the best way to learn short-form video is to immerse yourself in the feedback loop of having people watch your content and engage with it while you produce more.

You need to pick one tool, one idea and start posting. After enough interaction you start to “get” what makes things click.

Get a character, a piece of news and a popular song and mash them together into a whole that makes sense for your niche. Repeat that 100 times and find a style that works for you, and what works for your audience. Every copycat account that used “my style” became their own thing after a month or so and became successful in their own right.

Footnote On Tooling

Listing tools becomes outdated quickly. Keep in mind as you read that this list will most likely be out of date.

You need three things to get started: a linear video editor, an AI model capable of doing what you want, and footage.

There are many free editors you can use. CapCut is the default software for most short-form editors and its free version has everything you need to cut and move things around. You can also use it on your phone. For desktops, DaVinci Resolve is free and a professional-grade piece of software with a strong community of tutorials, plugins and even native integration for AI agents to use.

Additionally, if you have any coding agents, you can ask them to use ffmpeg to perform nearly any editing action you might need. The disadvantage here is that you won’t be able to jump in and perform minor edits yourself. Tools like Remotion are made to be “agent-first” and ship with a simple browser-based linear editor you can use. I do not recommend using these, however.

For AI models you need to decide if you’re willing to spend money or if you’re able to run things on your own hardware. For the sake of brevity, I’m gonna assume you’re willing to spend some money.

For dubbing characters there are two categories. General video models can do audio-driven video generation, matching the output to a song or voice line. If your goal is to dub a character singing or talking, any open-source model that has “audio-driven generation” will do the trick. You simply input a picture and your target audio and you’ll get the character performing. I recommend placing your target against a solid color matte background so you can more easily change what appears behind them.

The main drawback of these systems is that they will usually generate at most 15 seconds of footage at a time, and will require you to chain generations together. An alternative to that is dedicated puppet dubbing models that just move a character’s mouth to audio input for a minute or more.

For a basic brainrot video, here is a recipe:

Create a character with a solid color matte background, then feed that image to a video model of your choice with an audio file.

Tom the cat wearing a dark three-piece suit against a green-screen backgroundA grinning trollface character in a dark suit against a green-screen background
Characters on a solid color matte.
Animated Tom in a suit performing against a green-screen backgroundAnimated trollface character in a suit performing against a green-screen background
The animated characters.

After you have the characters moving, use a video editor to remove that color background with a “color key” and overlay that character on whatever background footage you want. This video netted me over 10 million views on TikTok and spawned a very popular trend in Eastern Europe.

Pink and purple stars sparkling against a black backgroundTom and the trollface character composited over sparkling stars, with a small Pleo mascot at the bottom
The finished composite. My old Pleo model was added to the bottom left in the video.

Example models you can use and their pricing

This is a short, non-exhaustive list of a few models you can check out. Most of the models here are open weights and you can run them locally if you have the necessary hardware. I have linked fal.ai as a provider because they have good prices and are easy to access but any provider would work. I am not affiliated with them in any way.

The table below refers to “general video generation”. You can use them to dub characters, generate footage from scratch or transfer a style to existing footage.

Pricing varies with model size, output quality, length, multi-modality (whether you are using reference audio, images or video) and providers. Keep in mind that this can get expensive quickly.

Model Provider 15 second clip
Wan 2.2 S2V (open weights) fal $1.50 to $3.00
Wan 2.2 S2V WaveSpeed $0.45
Wan 3.0 fal $0.75 to $3.00
Wan 3.0 Prime (faster) fal $1.02 to $4.20
Seedance 2.0 fal $4.55 to $10.23
Seedance 2.5 fal $3.31 to $7.10
LTX-2 Pro (audio-to-video) fal $1.50
LTX 2.5 Fast (audio-to-video) fal $1.95
LTX 2.5 Pro (audio-to-video) fal $2.55
MiniMax H3 (open weights) MiniMax $1.20 to $1.95
MiniMax H3 Max fal $0.60 to $1.20

Dedicated dubbing models

These models are tuned for dedicated avatar dubbing. They are concerned with generating longer clips of characters talking and not general footage generation.

Model Provider 15 second clip 60 second clip
Kling AI Avatar v2 fal $0.84 to $1.73 $3.37 to $6.90
Creatify Aurora fal, via Creatify $1.05 to $2.10 $4.20 to $8.40
VEED Fabric 1.0 fal $1.20 to $2.25 $4.80 to $9.00 (two generations)
InfiniteTalk fal $3.00 to $6.00 $12 to $24
InfiniteTalk kie $0.23 $0.90