How I Made a Professional, Cinematic Music Video in 2 Days Using AI

I’ve made many music videos. This is my most recent one. I released it on YouTube on July 20, 2026. I think it’s one of the best I’ve made because it’s more than just singing; it’s more like a story with cinematic shots throughout.

For reference, here’s the original music video, which was released on August 24, 2000.

Personally, I like my version of the music video more, not because I made it, but because the original video has seemingly irrelevant and random scenes, and I don’t care for some of the dance moves. Also, the actual singing only covers a portion of the vocal sections of the song, and the singing clips are often wide shots, so it’s unclear whether they’re actually singing or not.

Anyway, here’s how I made the music video.

1. Get a Song

Sometimes, I create music using AI in Suno. For this music video, I chose an existing Bollywood song called “Kya Maine Aaj Suno” from the movie “Harama Dil Aapke Paas Hai”. It’s an old song from August 24, 2000.

Note: I’m not Indian, and I don’t speak or understand Hindi. It’s just a song I had in my music library from a long time ago that I thought might make a good music video. I used Google Translate to translate the lyrics from Hindi to English, but the translation didn’t make much sense, so I used ChatGPT to make sense of it.

2. Create Subtitles

I used SubtitleEdit (free) to create the subtitles. SubtitleEdit can’t import mp3s, so I converted the song (mp3) to a video with a black screen in mp4 format and imported the video. I then manually created the subtitles since auto-subtitle generation is often wrong, especially for non-English audio. I made sure each lyric line time range started exactly or slightly before the vocals for that lyric.

3. Create Character Sheets

When I made music videos, I like to star in them, but I like to change my appearance, except my face, to match the theme of the song. So, for a Chinese song, I have AI create a character sheet of me but with typical Chinese clothing and a hairstyle suitable to the theme of the song. Here’s an actual image of me taken with my phone.

I then used ChatGPT to create a character sheet of me with a gold necklace. This is what it generated. You can start to see my facial identity drift a little, but it’s still close enough and acceptable.

I then told ChatGPT to replace my hat with hair of a male Bollywood singer and to add a full-body shot. This is what it created. You’ll notice that my facial details drifted even further, but I figured it was still close enough and acceptable, so I settled with this character sheet for the male singer (me).

For the female singer, I started with this photo.

I asked ChatGPT to remove the red dot on her forehead, remove her necklace and the gold thing on her head, and give her white, fitted pants. Here’s what ChatGPT generated. Since I wasn’t trying to recreate a real person, whatever looked good was acceptable, so I chose this character sheet.

4. Create Black-Screen Video Clips for Each Lyric

I used CapCut for video editing and SeeDance 2.0 to generate the final video clips for each lyric. SeeDance generates videos at 24 frames per second (fps), so I set CapCut to create 24 fps videos.

To have SeeDance generate an accurate lip-sync video, I need to give it a reference video with the singing audio. You can upload a reference audio to SeeDance, but for some reason it doesn’t work as well as a reference video. So, I just created a bunch of black-screen video clips, one for each video clip. SeeDance 2.0 supports video generation between 4 and 15s. I make sure each black-screen video clip starts when the vocals start for a particular lyric and ends on an integer number, e.g., 8s instead of 8s and 13 frames. Here’s what I did to create the black-screen video clips.

  1. Import the full song (mp3) to the timeline
  2. Extract the vocals using https://uvronline.app/ai (this is necessary for better lipsync by removing any background and instrumental sounds)
  3. Add the vocals audio (mp3) to the timeline
  4. Import the subtitles (srt file) into CapCut and add to the timeline

In order to see the duration of a clip, I found it easier to import a solid green image to CapCut, add it to the timeline. The left edge should line up with the start of the subtitle element. The right edge should end either at the end of the subtitle element or after it, and it should be a whole number in seconds, not seconds plus frames. I also like to add text to the timeline with just a number so I can see which video clip I’m working on. In the screenshot below, you’ll see the tracks from top to bottom are

  • text elements labeled 1, 2, and 3 to see which track I’m working on
  • solid green elements to easily see the start and end of a clip and the duration
  • the start and end of each lyric from the subtitles file
  • the vocals-only audio
  • the full-song audio

For clip 1, I decided to group lyrics 1 and 2 into one video clip. This particular song has alternating male/female vocals, e.g.,

  • lyric 1 is a female voice, which I prefixed with “F”, e.g., F Kya Maine Aaj Suna.
  • lyric 2 is a male voice, which I prefixed with “M”, e.g., M Haan Maine Tumko Chuna

Notice how the green element for this clip is exactly a whole number (9s long).

After repeating this process for each section that will be converted into a video clip, I

  • hid all visible elements (text and solid green),
  • disabled the full-audio track
  • enabled the vocals-only track

and exported each clip in 480p (the picture quality doesn’t matter since these videos clips are for audio reference only) and named each clip by its ID, e.g., audio1.mp4, audio2.mp4, etc.

For comparison, here’s clip 1, full audio and vocals only.

Note: sometimes, I would make a black-screen video longer than its lyric duration so its duration would be a whole number in seconds. I would then simply trim the end so that the clip ends when the next vocal segment begins.

5. Plan the Video Storyboard

Once I had all clips grouped in CapCut and all vocal segments exported as black-screen video clips, I created a storyboard in Excel like this. The yellow sections are instrumental sections. The blue sections are vocal sections. Since I wanted this music video to be more like a mini movie with a story as opposed to a bunch of random clips, I used this storyboard to help plan the story. I used ChatGPT to propose scenes for each clip.

View the Excel file

6. Create Lipsync Video Clips

Since lipsync video generation is difficult, I started with these clips and left the instrumental sections for last. I used SeeDance 2.0 via Kie.ai. At first, I tried SeeDance 2.0 mini, but the lipsync quality was bad and inconsistent. There’s SeeDance 2.0 Fast, but I decided to stick with the regular version of SeeDance 2.0. Since this version is expensive, I generated lip-sync clips at 480p and generated non-lip-sync (instrumental) clips at 1080p, since the higher the resolution, the higher the cost.

Whenever a clip included both characters, I added their character sheets as reference images. For the audio to be lip-synced to, I added the black-screen video clip. The duration is set to match the duration of the black-screen video. For the actual prompt, I asked ChatGPT to generate it for me. The prompts can become very long but very detailed, resulting in highly professional and cinematic results. For example, here’s one prompt for just one lip-sync clip.

Use the song from reference video 1 as the audio.

The characters must exactly match reference image 1 (male) and reference image 2 (female) throughout the entire video. Use the character sheets as the only source of truth for each character's identity, face, hairstyle, clothing, accessories, and overall appearance.

Exactly two people appear in the entire video: one male matching reference image 1 and one female matching reference image 2. No other people appear at any time.

Identity continuity is critical. From the very first frame to the final frame, the left character must always remain the male from reference image 1, and the right character must always remain the female from reference image 2. The female must already appear as the correct female in the very first frame, before she begins singing. Never duplicate the male. Never duplicate the female. Never swap, morph, replace, transform, or blend the two characters at any point. Changing singers must affect only lip-sync and facial performance. It must never change either character's identity, face, body, clothing, accessories, gender, or position in the frame.

Lyrics:

Tumko Pata Hai\nBolo Na Kya Hai

This is a romantic Bollywood duet.

The scene takes place on a beautiful tropical island wooden pier extending into crystal-clear turquoise water on a bright sunny afternoon. The luxurious white motor yacht from the previous scene is docked behind them along the pier, naturally continuing the story. Palm trees sway gently on the nearby white-sand beach while small tropical islands are visible in the distance beneath a brilliant blue sky.

The couple stands comfortably side-by-side on the wooden pier, facing slightly toward each other while enjoying the peaceful ocean surroundings.

There is absolutely no physical contact between them at any point. No hand holding, no touching, no hugging, no kissing, no leaning against each other, and no body contact of any kind. Maintain a small, natural gap between their bodies throughout the entire scene. Their romance is expressed entirely through warm smiles, affectionate eye contact, natural facial expressions, and relaxed body language.

The video is one continuous cinematic shot with no cuts.

The camera begins with a medium waist-up shot from slightly in front of the couple and slowly performs a smooth sideways dolly along the length of the pier while maintaining approximately the same distance from them. The movement should feel elegant, stable, and cinematic. Throughout the shot, the turquoise ocean remains visible on both sides of the pier while the yacht, palm trees, and tropical island scenery create beautiful depth in the background.

The female sings the lyric:

"Tumko Pata Hai."

She lip-syncs perfectly to the original audio while smiling playfully at the male, as though teasing him with a secret. Her expression conveys warmth, affection, and gentle curiosity.

As she finishes, the male immediately replies:

"Bolo Na Kya Hai."

He lip-syncs perfectly to the original audio while smiling warmly back at her. His expression conveys playful curiosity, happiness, and affectionate encouragement, inviting her to continue. He may briefly raise his eyebrows with a natural friendly expression before smiling again.

Only the currently singing character lip-syncs the lyrics. The other character maintains a gentle, natural smile with subtle facial movements, breathing, blinking, and realistic expressions, but never mouths the lyrics or appears to sing. Changing the active singer must never change either character's identity or appearance.

Both characters blink naturally, smile subtly, and make small natural head movements. Their body language should feel relaxed, elegant, affectionate, and comfortable together while always maintaining the small gap between them. Avoid exaggerated acting or large gestures.

A gentle tropical breeze softly moves their hair and clothing. Bright sunlight creates sparkling reflections across the turquoise water and soft natural highlights on their faces, giving the scene a luxurious, peaceful, and romantic atmosphere.

No scene changes. No cuts. No dancing. No text. No subtitles.

Modern Bollywood movie style. Bright tropical daylight. Crystal-clear turquoise water. Rich saturated colors. Highly realistic. Beautiful cinematic composition. Natural facial expressions. Accurate lip-sync that follows the original reference audio exactly.

As an example, here’s one generated lip-sync clip. Notice the audio is the vocals-only version. Later, this audio from this clip will be muted and replaced with the full audio.

7. Create Additional Reference Images

When a series of video clips contains the same object and object consistency is important, I create a reference image for that object. For example, in the instrumental intro of the music video, I show a Rolls-Royce convertible across multiple clips. To ensure SeeDance creates a consistent-looking car, I created this reference image.

In addition to the two character reference images, I added this car image as a third reference.

8. Review, Export and Upscale Final Video

Once all clips have been generated and added to the timeline, I previewed the entire compilation and when it looked good, I exported it at 720p since my instrumental clips were generated at 720p. The 480p lip-sync clips would just get upscaled to 720p, but not using AI. I then reviewed the 720p-version of the full video and when it looked good, I upscaled it to 4K using Topaz Video AI. If you want the best quality video, you can have SeeDance generate video clips in 4K, but it will be much more expensive than 480p and 720p.

9. Generate YouTube Thumbnail

To generate the YouTube thumbnail image, I used SeeDream 5.0 on Kie.ai. ChatGPT gave me the prompt and I uploaded the two character sheets. Here’s the generated thumbnail that I approved.

How I Converted a Half-Bathroom into a Full-Bathroom

I got a permit to convert a half-bathroom into a full bathroom. The half-bathroom actually had a 2.7’x2.7′ shower in it, but it wasn’t permitted and was too tiny to shower in. I also didn’t like that the shower was raised above the floor because the previous owner didn’t want to bury the drain pipe in the concrete slab.

Getting the permit was easy since the permit type was “OTC” (Over The Counter), which gets approved in a day instead of months. I didn’t need perfect code-compliant technical drawings. I was allowed to just sketch the changes I was proposing. I decided to use a 2D CAD web app called Rayon.

In order to install a 3’x4′ shower, I needed to move the furnace return ducting out of the way. An HVAC technician ended up moving it to the ceiling in the hallway. This required a separate permit, a mechanical permit. I ended up paying about $600 for the bathroom remodel permit and $400 for the permit to move the furnace return duct. Here are photos of the bathroom remodel.

Before the remodel, showing the old, tiny, useless shower
The furnace. The return duct was below the furnace.
After removing the furnace, you could see the opening where return air flow passed.
After removing the furnace stand, you could see the opening in the wall where return air flowed through. I needed to move this duct to make space for a larger shower.
To demolish the old shower, I used a jackhammer with a chisel bit.
With the walls removed, you can see the return air filter in the hallway wall, which later was moved to the hallway ceiling.
Plumbing for sink. On the left is the drain pipe for the washing machine in the garage.
I wanted the new shower to be level with the floor instead of raised, so I needed to cut the concrete slab to run the drain pipe in it. The shower drain would tap into the existing sink drain. I marked the slab where I needed to make cuts using a saw.
After cutting the concrete with the saw, I used a jackhammer to break the concrete.
The saw cuts prevent the jackhammer from breaking concrete outside the cuts.
I used an angle grinder to cut rebar in the concrete.
We cut the old copper drain pipe with an angle grinder so we could put a new ABS pipe with a tee for the shower drain pipe.
Since the width of the bathroom (4′ 7″)was longer than the length of the shower pan (4′), I had to fur out the wall by adding two 2x3s. This worked in my favor because it provided space to install the shower supply line plumbing.
The new ABS shower drain pipe sloped a bit for drainage and was temporarily secured to the rebar using zip ties. I wrapped the pipes in foam so if I later need to cut the concrete to work on the pipe, the concrete won’t stick to it. Since I was going to add a kitchenette in the garage for the JADU, I added supply lines for it so I wouldn’t have to reopen the bathroom wall to add them later.
We used rubber couplings with worm clamps to connect the old and new pipes.
For renter convenience, I installed a bunch of shower nooks. I should have installed the metal kind that doesn’t require tiling, because the effort and cost of tiling these nooks were high.
Per code, I had to separate the light switch from the exhaust fan switch.
The light switch has to turn off automatically when no motion is detected and the exhaust fan has to turn on and off automatically when moisture is/isn’t detected. California building codes are SO PICKY!!!
I covered up the hole in the wall where the furnace return air flowed.
I also built a new furnace stand using sturdy metal brackets.
I installed mold-resistant drywall, but the inspector said I had to replace it with Denshield.
The HVAC technician reinstalled the furnace on my new stand and opened the side panel of the furnace for the return air to flow through the side instead of the bottom.
The HVAC technicians cut insulation and metal sheets to build the new return air duct.
The furnace return air now flows into the ceiling…
I also put plastic to serve as a vapor barrier and to make concrete removal easy in case I ever need to cut the concrete to work on the drain pipe.
I then poured concrete in the trench.
Before installing the shower pan, we put a layer of plastic and a bed of concrete. The plastic would make it easy to remove the concrete in case we make a mistake. The concrete serves as a solid foundation for the hollow shower pan.
I replaced the purple mold-resistant drywall with Denshield as required by the inspector.
To make the top surface of the trench smooth with the existing concrete foundation, I poured a bit of cement (not concrete) and smoothed it using a rotary sander.
Before installing tile, I placed a sheet of drywall on the shower pan so if a tile falls, it wouldn’t damage the pan.
I hired a tile contractor to install tile in the bathroom. Here, he’s shown mixing mortar.
The tile contractor used a laser level and started placing mortar on the wall.
For straight tile cuts, he used a manual tile cutter. Otherwise, he used a wet tile saw.
He used spacers between tiles and screw-on levelers to ensure adjacent tiles are level.
For the baseboard, I had the contractor cut floor tile into 3″-wide x 24″-long strips instead of buying bullnose tile. I then had him install metal edge trim.
The last step was to put grout in the joints.
Overall, I was very happy with the results, including the tile and grout color.
However, to save money and time, and for a better appearance, I should have just bought the no-tile metal shower nooks like the black metal footrest.
I installed a sliding shower door.
The baseboard came out really nice as well.
Here’s how it looked after reinstalling the toilet and sink.
For the sink drain and p-trap, I used a flexible drain. Much easier to install and more reliable than the rigid plastic ones.

WatchWise: Watch YouTube More Efficiently

A free Chrome extension that helps you decide if a video is worth watching, generate summaries, and jump directly to the topics that matter.

Every day I watch YouTube to learn something.

Programming.
AI.
Home improvement.
Real estate.
Finance.

The problem isn’t finding videos.

It’s figuring out whether a 45-minute video is actually worth watching.

Too often I click a promising title only to discover:

  • the answer could have been explained in 3 minutes,
  • the title exaggerates what the video actually delivers,
  • half the video is filler.

After wasting enough hours on videos like this, I decided to build a small tool, a Google Chrome extension, myself.

I call it WatchWise.


What WatchWise Does

WatchWise adds a small, red WW button to YouTube, located in the top-right corner.

One click expands the WatchWise panel, giving you three built-in tools.

1. Title vs Content

Instead of guessing whether a video delivers on its title, WatchWise asks YouTube’s built-in AI and formats the result into an easy-to-read report.

It tells you things like:

  • Does the content actually match the title?
  • How much of the video stays on topic?
  • How much time is spent on filler?

Here’s a screenshot explaining why a particular video is mostly clickbait and a waste of time to watch.


2. Summary + Table of Contents

Sometimes you don’t need to watch the entire video.

WatchWise generates:

  • a concise summary
  • an organized table of contents
  • bulleted list of key takeaways for each section
  • clickable timestamps

You can immediately jump to the section that interests you.


3. Smart Chapters

For long interviews and podcasts, WatchWise organizes the discussion into structured sections that read almost like an article.

Instead of scrubbing through a one-hour conversation, you can browse topics and decide where to start.


Why I Built It

This isn’t another AI chatbot.

It simply makes YouTube’s existing “Ask about this video” feature much easier to use.

Instead of writing prompts every time, I click one button.

That’s it.


Privacy

WatchWise:

  • doesn’t require an account
  • doesn’t collect personal information
  • doesn’t use analytics
  • doesn’t send data to my own servers

Everything happens inside your browser using YouTube’s existing Ask AI feature.


Try It

You can install WatchWise free from the Chrome Web Store.

Chrome Web Store:
https://chromewebstore.google.com/detail/watchwise/baclglojkadoiopgnjakdajnjlbcgekc


Final Thoughts

I originally built WatchWise for myself because I was tired of wasting time on clickbait and overly long YouTube videos. I also wanted a short executive summary in bulleted list format with clickable timestamps as well as a longer summary.

If it saves other people time too, then it has already accomplished its goal.

Easily Generate and Edit Subtitles or Lyrics From a Video With Subtitle Edit

Subtitle Edit is a free app that lets you generate and edit subtitles from a video. If you have a song that you want the lyrics for, you can export the song as a video and then add the video to Subtitle Edit. In the example below, I want the lyrics for an Italian music video.

Download Subtitle Edit and then install it

Click Video > Open video… and select your video.

Click Speech to text…

In the pop-up window, I like to choose

  • Engine: Whisper CPP
  • Model: large-v3-turbo (1.5 GB)

Since my video is in Italian, I set the language option to Italian.

When processing is done, you will see the lyrics in the left pane. Clicking on a verse will jump the playhead to the point in the waveform where that verse begins. The “Text” field lets you edit subtitle text. In the waveform, you can also drag the vertical start and end lines for each verse, which will update the timestamp accordingly.

Delicious Air-Fried Chicken Breast With Tzatziki

Everyone knows that the problem with baking or air-frying chicken breast is that it often comes out too dry. Most people prefer chicken breast over other parts (legs, thighs) because it contains less fat, but it’s the fat content that makes chicken legs and thighs more juicy. You can dip chicken breast in barbecue sauce, but many such sauces aren’t healthy. One option that I found to work really well is Tzatziki. Not only is Tzatziki very healthy, but it also makes air-fried chicken breast, including dry ones, tasty and moist. Here’s my recipe for making this.

Ingredients

Instructions

  1. Slice the chicken breast in half (I like to slice 95% of it so it’s still connected)
  2. Massage oil on both sides
  3. Sprinkle seasoning on both sides
  4. Air fry it till the internal temperature reaches 165 degrees. (In my t-fal air fryer, that’s about 13 minutes, flipping half way)
  5. Remove from air fryer
  6. Using a spoon, put Tzatziki over the top half until it’s mostly covered.
  7. Enjoy

Photos

Alternatively, for a lower- calorie option that still tastes good, replace Tzatziki with avocado salsa.

Generate Consistent Characters in Videos with SeeDance 2.0

Most AI image-to-video generation tools support first-frame reference images. Considering how much more expensive video generation is compared to image generation, it makes sense to use image references, like a first frame, when generating videos. However, providing a first-frame image, with or without a last-frame reference, still fails when you need character consistency because the AI model only knows what the character looks like in the first-frame image. Fortunately, SeeDance 2.0 supports multiple reference images, so you can upload both a first-frame image and a character sheet containing different views of a character.

For example, I had the following character sheet.

I then used it to create the following first-frame image using Nano Banana 2 in OpenArt.ai.

If I zoom in, I can see that the face looks close enough to the one in the character sheet.

When I created a video using Kling 2.5 of the woman walking, using that image as the first frame, I got the following.

The video starts out fine because of the first-frame reference, but as it progresses, the woman’s face slowly changes and looks less and less like the one in the character sheet. Here’s a screenshot of just her face in one frame of the video.

What’s particularly different is the nose, but the width and height of her face looks somewhat different as well, especially compared to the character sheet.

Now, let’s see how the same video turned out using SeeDance 2.0 with multiple references. For this, I used Kie.ai.

Since I wanted to keep the setting and just replace the subject, I used Photoshop to “remove” the subject from the previous image. I selected the subject and clicked “Remove”, which used AI to remove the woman.

This is what I got.

Next, I upload the character sheet and setting image to Kie.ai (SeeDance 2.0 page), gave it the same prompt I used in Kling 2.5.

Here’s the resulting video.

Notice how the character looks EXACTLY like the one in the character sheet throughout the entire video clip.

Here’s a close-up of the face near the end of the clip.

My Favorite Oscillating Tool Attachments

HERCULES Multipurpose Drywall Blade for Oscillating Multi-Tools

Purpose: Cutting drywall

HERCULES 2 in., 3-in-1 Multifunction Knife Blade for Oscillating Multi-Tools

Purpose: Cutting carpet, cardboard, roof shingles

HERCULES 1-1/4 in. Tapered Scraper Blade for Oscillating Multi-Tools

Purpose: Removing caulk and sealant

HERCULES 1-3/8 in. High-Carbon Steel Reduced Neck Plunge Blade for Oscillating Multi-Tools

Purpose: cutting wood

This particular blade worked impressively well when I used it for cutting nails in 2×4 studs and for cutting notches in 2x4s.

HERCULES Triangle Sanding Backing Pad for Oscillating Multi-Tools

Purpose: sanding, including sanding drywall

HERCULES 3-3/8 in. and 4-1/4 in. Recessed Light Cut-Out Saw Set for Oscillating Multi-Tools, 2-Piece

Purpose: cutting holes in walls, including for recessed lights

And here’s an oscillating tool that I use. It’s low-noise, low-vibration, and has a trigger-lock feature, tool-free accessory change, and a light.

Make Realistic Lip-sync Music Videos with SeeDance 2.0

I just made this music video, and the lip-sync portion is amazingly impressive.

I actually used SeeDance 2.0 Fast at Kie.ai, but you can use SeeDance 2.0 as well and get up to 1080p resolution. For each generation, I used

  • Prompt
  • Reference image (not first-frame image)
  • Reference video (this was just a black video containing the audio clip)
  • “Generate audio” enabled
  • “Web search” disabled
  • Duration = duration of reference video

SeeDance 2.0 supports generating videos up to 15 seconds long. But, if you give it a 15s reference video and you want to lipsync a character in it, the lipsync won’t work. So, when generating lipsync videos, always provide a reference video that is no longer than 13 seconds to be safe.

When creating reference videos, make sure the duration is a whole number, not a fraction, e.g., 5 full seconds, not 5.5 seconds. The reason is because in the UI, Kie.ai or another app may round down the duration to the nearest whole second, and if you tell SeeDance you want to generate a 5s video, then it will generate a 5s video, not a 5.5-second video, and your lip-sync video will be truncated. I use Capcut to generate my black reference videos. I put a playhead at a location where I want a segment to begin and end and set a marker at each location, making sure the time ends with :00 (no frames), e.g.

  • start 2:54:00
  • end: 3:04:00
  • duration = 10s

If I really need to split at a location between seconds, like 2:54:09, then make sure the end location includes the same number of frames, e.g., 3:04:09, so you end up with a duration in whole seconds.

SeeDance 2 supports reference audio, but for some reason, it didn’t lip-sync my reference image, and sometimes it would change the lyrics.

Also, the following method worked well for English audio. It may not work for other languages. If you find that it doesn’t work for your language, then see some options below.

Here’s a screenshot of the inputs.

Below are the inputs and outputs for various lip-sync clips.

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

Reference image:

Reference Video:

Output:


Prompt:

use the song from the given video and use the character from the given image to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing.

A man (same subject, unchanged face and outfit) singing into a microphone at the center of a large ancient Roman-style amphitheater at night. Camera is positioned at chest height, medium close-up framing, stable and focused on the singer.

The audience fills the stone bleachers behind him, hundreds of people seated and standing, naturally animated: subtle head movements, clapping, cheering, shifting in seats, occasional phone screens glowing, realistic variation in motion without repetition.

Warm golden stage lighting illuminates the singer from the front and slightly below, creating a cinematic glow on his face. Behind the singer, rows of soft amber lights line the steps and columns. Moving stage lights sweep slowly across the audience and architecture, creating gentle light motion across the crowd and stone surfaces.

The night sky is clear with visible stars. Light atmospheric haze adds depth and catches the beams of moving lights. The columns and amphitheater remain stable and realistic.

The singer performs naturally: subtle head movement, mouth lip-syncing accurately, slight body sway, breathing and posture shifts.

Camera behavior: very subtle cinematic push-in (slow, minimal), no drifting or unintended orbit, no zoom jitter. Maintain subject as the clear focal point at all times.

Depth of field: subject sharp, audience slightly softened but still readable.

Lighting style: warm amber/yellow tones only, no harsh white light, no overexposure, cinematic contrast.

Reference image:

Reference Video:

Output:

Here are some similar clips using the same prompt and reference image but different reference videos (for the audio).

Reference Video:

Output:

Reference Video:

Output:

Playing Musical Instruments That Sync to Music

SeeDance 2.0 also seems to support making a video of a person playing a musical instrument in a way that matches the sounds in a reference source. Consider the following:

Prompt:

use the song from the given video and use the character from the given image to make a music video of the man playing the guitar sounds in the song. sync the playing of the guitar to the guitar sounds in the song.

Reference image:

Reference Video:

Output:

Here’s another example.

Prompt: use the song from the given video and use the character from the given image to make a music video of the man playing the saxophone such that the sound of the saxophone in the song is in sync with the playing of the saxophone. the background should be solid green as in the reference image for chroma key background removal later on. the camera remains fixed. do not zoom in or out. the man’s fingers move on the saxophone naturally and in sync with the sound from the song. the man moves his body naturally as he plays the saxophone.

Reference Image: https://www.youtube.com/watch?v=_piqiZkLKgY

Reference Video:

generate_audio: on

The duration was set to the duration of the reference video. I used SeeDance 2.0 Fast and 720p resolution. I later upscaled the video to 4K using Topaz Video AI.

Result

SeeDance 2.0 Error – output audio may contain sensitive information

If you get an error that says “The request failed because the output audio may contain sensitive information.”, then disable audio generation.

For example, in order to make the following video,

I had to use the following settings in Kie.ai:

Prompt: use the song from the given video and use the character from the given image to make a music video of the man singing the song in front of a green screen, as shown in the reference image. he stands in place and sings the exact lyrics in the song audio as if lip-syncing to the audio with natural face and body movements, but keep his hands beside his body. The camera is fixed and doesn’t zoom in or out and doesn’t pan.

Do not add shadows, floor shadows, lighting gradients, reflections, stage lighting, environmental lighting, or any background elements. Camera locked off and completely static.

Reference Image:

Reference Video: (example)

generate_audio: off

The duration was set to the duration of the reference video. The resulting clip was

I then removed the green background in Capcut to overlay the singer on a series of background video clips.

Singing Lip Sync Videos Using HeyGen

If your song is not in English and SeeDance 2.0 can’t lip-sync it correctly, then use HyeGen with custom motion enabled, as follows.

Log in to HeyGen and create an avatar. You can simply upload a photo of your singer. I used the one below. I put my avatar on a green background so I can chroma key it out.

Open Avatar Studio and

  • in the Script section on the left, instead of typing your script, upload your song’s audio (mp3)
  • in the Avatar and Voice section on the right, under Voice, you can ignore this since you’ll be using the audio you uploaded
  • in the Avatar and Voice section on the right, under Motion Engine, choose “Avatar IV”

then, and this is important, click the “Advanced Settings” button.

Toggle on “More expressive motion” and enter a custom motion prompt.

Optionally, you may click the “Generate motion prompts” icon, which will generate motion tags as shown below.

Then, click the Generate button.

Following are examples comparing different settings.

HeyGen LipSync Using Avatar IV WITHOUT Custom Motion

HeyGen LipSync Using Avatar IV WITH Custom Motion

In this example, I didn’t click the “Generate motion prompt” button.

HeyGen LipSync Using Avatar IV WITH Custom Motion

In this example, I did click the “Generate motion prompt” button.

As you can see, in the first example, the avatar doesn’t look like he is singing at all, and in the last 2 examples, the avatar looks more expressive. It may be difficult to tell the difference for such a short clip, but the difference is actually huge when you lipsync a full song, as in the following example.

The lip-sync quality is definitely not as good as SeeDance 2.0, but it seems to be the best option when SeeDance 2.0 doesn’t work for a particular language.

UPDATE 6/5/2026

There’s another way to generate lip-sync videos using SeeDance 2.0, and it supports non-English languages. Here, I’m using Kie.ai. Instead of uploading a black video with audio, I upload an audio and include the lyrics in the prompt.

Inputs:

Prompt: Lyrics: Naik bajaj jingga bunyinya setengah mati

The guy in reference @image1 sings the verse in @audio1 in a music video way. The verse in the lyrics is in the Indonesian language. Keep the camera fixed. Don’t zoom in or out. Keep the background solid green as in reference @image1. The man in reference @image1 moves his body naturally in a music video way.

Reference Image:

Reference Audio:

Duration: 6s

Output:

Inputs:

Prompt: Lyrics: Naik bajaj jingga bunyinya setengah mati

The guy in reference @image1 sings the verse in @audio1 in a music video way. The verse in the lyrics is in the Indonesian language. Keep the camera fixed. Don’t zoom in or out. Keep the background solid green as in reference @image1. The man in reference @image1 moves his body naturally in a music video way.

Reference Image:

Reference Audio:

Duration: 6s

Output:

UPDATE 6/13/2026 – Actually, using a video reference containing the audio is better than an audio reference. See following example.

Inputs:

Prompt: Lyrics:

Còn tôi như cánh chim
Sẽ bay đi muôn phương
Mang về mầm xanh tươi

use the song from the given video (@video1) and use the character from the given image (@image1) to make a music video of the man singing the song. he sings the exact lyrics in the song as if lip-syncing. The lyrics are in Vietnamese. He sings passionately and moves his body naturally to the sound of the music. Keep the camera fixed. Don’t zoom in or out. Keep the background solid green as in reference @image1.

Reference Image:

Reference Video:

Duration: 11s

Output:

Create Cinematic, Multi-Shot Lip-sync Music Videos

To create cinematic, multi-shot lip-sync music videos in one SeeDance 2.0 video generation, do the following:

  1. Give Claude or ChatGPT the lyrics to the whole song so it knows what the song is about
  2. Create reference video clips in 720p containing audio segments that are 14s or less. Don’t split mid-word.
  3. For each clip, give Claude the mp3 and the lyrics for that clip, if any, and tell Claude you want a SeeDance prompt to generate a music video. Specifically, tell Claude to give you the shots (scenes) similar to the example below.

Shot 1: Medium-close on the singer at golden hour along a cliffside coast, glowing amber coastline and ocean curving behind him, warm sun on his face. Camera slow gentle push-in. He is the only person in frame.

Shot 2: Medium shot of the singer standing at a coastal overlook, vast golden California coastline stretching into the distance behind him, soft waves and warm haze. Camera slow drift. He is the only person in frame.

Shot 3: Medium-close, front-on, on the singer with the blazing golden sunset coastline glowing behind him, the warmest light of the clip full on his face, a peaceful contented expression. Camera slow push-in. He is the only person in frame.

Then, append it to your base prompt, which is

LYRICS: “[enter lyrics for the clip / segment here]”

use the song from the given video (@video1) and use the character from the given image (@image1) to make a music video of the man singing the song

@image1 is the face and identity reference for the lead singer — match his face, afro, beard, and glasses to @image2 throughout, keeping his identity consistent.

The generated audio must match the audio in @video1 EXACTLY and the lip sync must match the vocal segments in @video1 EXACTLY.

4. Add your character sheet as the first reference image (@image1).

5. Add your reference video

6. Specify a duration that matches the reference video duration

Example Character Sheet Image

Example Reference Video

Output

Wall Framing: How to Cut 2×4 Wood Studs So They Fit Perfectly Between Top and Bottom Plates

This method assumes the top and bottom plates are already in place.

  1. Use a laser measure like the one shown below. In this example, the reading shows 8′ 11″ 5/16″. If you prefer precision to 1/8 of an inch, you may be able to change the precision settings in your laser measure.

Make sure the laser dot is on the surface of the distance you want to measure.

2. Then, using a measure, mark that length on your 2×4 wood.

3. Use a carpenter square to draw a straight line at that mark

4. Cut the wood at that line using a miter saw.

If you want to dry-fit your studs before permanently fastening them with nails, you can use A34 metal brackets with screws.

Chia Seed Pudding Recipe

Ingredients

  • 2 tbsp black chia seeds
  • 8 oz of full-fat coconut milk
  • 1.5 tbsp monk-fruit sweetener
  • 1/4 tsp vanilla extract
  • Fruit topping (optional)

For the coconut milk, make sure to get full-fat coconut milk that’s usually in a can, not the heavily diluted coconut milk that’s in a carton, unless you want less calories at the expense of taste.

Instructions

  1. Add and mix all ingredients in a container
  2. Place in refrigerator for 1 hour
  3. Mix again to break up any chia seed clumps
  4. Refrigerate again for another 1-2 hours or overnight
  5. Chia seeds should have absorbed the milk and become soft
  6. Optionally, top with fruit (blueberries, raspberries, etc)
  7. Enjoy