Bidirectional SAM2 Tracking for Video Inpainting in ComfyUI
Bidirectional SAM2 Tracking for Video Inpainting in ComfyUI â on a 16GB Laptop GPU
A local, open-source pipeline for replacing a subjectâs head in a moving video shot, and the SAM2 patch that makes bidirectional mask tracking possible.
A few months ago I published an open-source project on GitHub: ComfyUI-Bidirectional-SAM2-Inpaint. It is an automated video inpainting pipeline built on ComfyUI,âŠ
This Saturday Iâll be talking about Liquid Horizons â the immersive four-wall panoramic installation currently on view at f|x|r gallery â and the process behind it: AI-generated imagery, projection, and the shifting horizon between the real and the synthetic image.
Canât make it in person? Step into the virtual gallery here:
Liquid Horizons â Immersive virtual gallery experience by ivo3d.com
âAmong closed-source models, I have worked most extensively with Googleâs ecosystem, hence it will remain a focal point of this analysis.
âStrategic Positioning & Core Functionality
While not the optimal choice for purely artistic image generation, recent developmental shifts indicate this is no longer their primary objective. The nanobanana2 model essentially positions itself as an AI-driven equivalent to Photoshop, and analogously, the latest Omni video model functions as an After Effects for video synthesis.
âImage Editing Workflows & Latent Space Degradation
In the realm of image editing, the platform has become indispensable. However, workflow optimization dictates editing one element at a time, as the system tends to lose track of complex, multi-layered prompts. Its built-in editor is highly viable for this purpose, despite the frustrating necessity of repeatedly applying negative prompts to exclude unwanted artifacts from the final output.
âSurprisingly for an autoregressive model, it exhibits a characteristic degradation in image quality with each successive edit, indicating that the entire image is being forced back through the latent space. Consequently, best practice requires archiving individual generation phases and compositing them externally in Photoshop.
âPrompt Architecture & Aesthetic Averages
Absent a highly complex prompt, the output regresses heavily toward the statistical average. Therefore, it is not the recommended starting point for artistic conceptualization; however, its advanced spatial awareness makes it highly adept at subsequent structural editing. Avoiding the default "Veo aesthetic" demands extremely detailed prompt engineering with strong stylistic descriptors. This exact requirement becomes a liability during debugging, as high prompt density makes it difficult to isolate semantic misinterpretations.
âAlignment, Safety Guardrails, and Open-Source Relevance
The platform's stringent safety censorship frequently causes workflow frictionâa factor that inadvertently preserves the market relevance of open-source alternatives. This is particularly evident during video editing with the Omni model within Google Flow. (For image generation as well, utilizing Flow is highly recommended over the standard Gemini consumer interfaces).
âThe corporate alignment appears hyper-focused on deepfake prevention; for instance, the model consistently refused to perform a simple head-swap on a masked figure. This necessitated authoring a custom workflow in ComfyUI, a topic slated for future discussion.
âVideo Synthesis & Physical Simulations
âTemporal Consistency: Despite its autoregressive architecture, it occasionally drops character-specific attributes during dynamic motion, though it still maintains temporal consistency better than its competitors.
âAbstraction vs. Physics: The model demonstrates a severe deficit in visual abstraction and classic animation principles. However, this is a current industry-wide limitation across all models, providing human artists with a continued competitive buffer. It compulsively defaults to cheap particle/glitter effects when tasked with abstraction, though it conversely exhibits a solid grasp of physical systems like rigid body dynamics and fluid simulations.
âLanguage Parameters: For video synthesis, English remains the optimal command language.
âConclusion & Current Bottlenecks
Strategically, the primary use-case appears to be business presentation asset generation rather than cinematic art. Despite the PR-driven exaggerations of official demos (a discrepancy the platform itself practically acknowledges), iterative prompt engineering significantly improves the execution rate of complex tasks.
âKey technical limitations persist:
âVideo extensions remain frustratingly plagued by temporal jumps, though these can now be cleanly edited out in post-production.
âThe progressive degradation of quality is likely a persistent technological artifact that will eventually require switching models or an architectural paradigm shift.
âThe absence of native 4K resolution remains a glaring vulnerability.
Since I was asked to use ai generators. (Disclaimer:)
Going backward in chronological order, Iâll start the mini AI series with the latest oneâthe very last one Iâve been tinkering with for just three months is the Grok model. I had to use it for images because Google censors them senselessly; for example, an image of Adam and Eve under the Tree of Knowledgeâwell, that just wonât work if itâs censored. But more on Google later. Grok feels like a model similar to Flux; it handles fashion and film noir very well, and itâs even got the Tumblr style down too. Good at vector styles. I like that itâs bold and has its own style/grok feeling. It tends to lean toward manga, and its anatomy isnât quite up to the standard youâd expect from a flagship model, but you can actually use that to your advantageâitâs better at morphing than most others out there today, and itâs good at editing too, though you can only do that with prompts. Itâs constantly being developed; the app updates daily, video editing abilities are constantly updated, andâsurprisinglyâit is still handles soft adult content, making it usable for art-related topics because of these features. So thereâs no need to look down on it. Thereâs a difference in quality between generated images and videos, but thatâs true of all models; the 30-second extension and max 1080p resolution are starting to feel a bit limited these days, but itâs still fun to use. You still have to deal with censorship, though. Secretly, Iâm hoping X will scrap the whole thing and Claude will get it as an image and video generator :) I was able to make a video with it that you wouldnât even guess was AI.
ki â KĂŒnstliche Intelligenz â das Lustige daran ist, dass dies mein ungarisches Monogramm ist. (Das Logo im Grok-Stil â eigentlich habe ich es ein wenig von Hand in Photopea bearbeitet, aber dabei die Logik durcheinandergebracht â also muss ich noch weiter daran arbeiten)
So, the thing is, Iâve been spending a ton of time again (sometimes 12 hours a day, working on multiple models simultaneously) tinkering with the AI, and itâs about time I wrote a proper summary of my experiences. But since this involved the entire pipeline, Iâll need to break it down into topics. First, I should write about image manipulation, simply because thatâs where I started 26 years agoâand I did that for 12 hours a day for a year and a halfâand back then, it was exclusively Photoshop. So Iâll try to set aside some time for that too, because, after all, we should only form opinions about what weâve actually learned.
Project Description: "A Plague in London" is a mixed-media production situated at the intersection of performative arts and digital scenography. Based on Daniel Defoeâs 1722 archival-style account, the work explores the structural and psychological dynamics of systemic crisis and social isolation. Directed by KristĂłf SzabĂł, the production utilizes Defoeâs text as a conceptual framework to analyze the sociopolitical mechanisms of a pandemicâranging from institutional denial and the enforcement of urban confinement to the eventual erosion of communal structures.
Visual and Medial Strategy: The production is characterized by a multi-layered visual dramaturgy, where IvĂł KovĂĄcsâs video art functions as a spatial and temporal mediator. Rather than serving as a purely illustrative background, the digital layer operates as a generative environment that reflects the interiority of the protagonists and the oppressive atmospheric silence of a city under lockdown. The visual components analyze the tension between the historical archival source and contemporary digital aesthetics, translating the 17th-century experience of quarantine into a modern medial discourse. KovĂĄcsâs work focuses on the abstraction of architectural boundaries and the visualization of the "Planet of Viruses" concept, examining the human condition through a lens of digital depth and spatial deconstruction.
Credits:
Video Art: IvĂł KovĂĄcs
Artistic Direction, Set Design & Sound Dramaturgy: KristĂłf SzabĂł
Performance: Maximilian von MĂŒhlen, Boshi Nawa, Lili Oksanen
Source Material: Daniel Defoe: A Journal of the Plague Year (1722)
Project Description: "A Plague in London" is a mixed-media production situated at the intersection of performative arts and digital scenography. Based on Daniel Defoeâs 1722 archival-style account, the work explores the structural and psychological dynamics of systemic crisis and social isolation. Directed by KristĂłf SzabĂł, the production utilizes Defoeâs text as a conceptual framework to analyze the sociopolitical mechanisms of a pandemicâranging from institutional denial and the enforcement of urban confinement to the eventual erosion of communal structures.
Visual and Medial Strategy: The production is characterized by a multi-layered visual dramaturgy, where IvĂł KovĂĄcsâs video art functions as a spatial and temporal mediator. Rather than serving as a purely illustrative background, the digital layer operates as a generative environment that reflects the interiority of the protagonists and the oppressive atmospheric silence of a city under lockdown. The visual components analyze the tension between the historical archival source and contemporary digital aesthetics, translating the 17th-century experience of quarantine into a modern medial discourse. KovĂĄcsâs work focuses on the abstraction of architectural boundaries and the visualization of the "Planet of Viruses" concept, examining the human condition through a lens of digital depth and spatial deconstruction.
Credits:
Video Art: IvĂł KovĂĄcs
Artistic Direction, Set Design & Sound Dramaturgy: KristĂłf SzabĂł
Performance: Maximilian von MĂŒhlen, Boshi Nawa, Lili Oksanen
Source Material: Daniel Defoe: A Journal of the Plague Year (1722)