Jump to main content

Hotel Lobby AI Prompt: The One We Send, Plus 5 to Copy

A Hotel Lobby AI prompt has to pin down the orange set, the one shared mic, who stands where and who raps. Plenty of prompts online are guesswork; the two below are the real ones our studio sends for every take. The first runs on Wan 3.0, our default, and keeps the Hotel Lobby track. The second is the Seedance version, which writes a fresh rap. After them come five prompts we wrote for tools that accept only two photos, then a line-by-line explanation and a fix-it table. Or skip all of it: the studio adds the prompt for you.

Last revised

A notebook of directions beside reference portraits and a stage photo.

HOTEL LOBBY AI PROMPT

Direct the performance

  • People
  • Motion
  • Sound
AI · Illustration · Hotel Lobby Video

The Wan 3.0 prompt behind Hotel Lobby Video

This version serves two photos in any frame shape; the aspect ratio travels as a model setting, not as words. Before the take starts, GPT Image 2 redraws each photo as a full-length studio portrait of the same person or pet. The model then gets four attachments, always in this order: Image 1, the left portrait; Image 2, the right portrait; Video 1, the Hotel Lobby reference clip with its sound removed; and Audio 1, the Hotel Lobby track trimmed to the same length.

Hotel Lobby AI prompt (Wan 3.0)
Follow Video 1 from start to finish: copy its framing, its timing and every movement, gesture, head turn, facial expression and mouth shape, along with the plain orange background and the silver microphone hanging from the ceiling between the two performers. Keep the camera framing of Video 1: no push-in, no pull-back, no pan. Do not add any movement or scene element that is not in Video 1.

Use Audio 1 as the background music for the whole video, exactly as it is: it is the entire sound of the video. Do not write new lyrics, do not re-sing it, do not add other music or voices.

The performer on the left is the subject in Image 1; the performer on the right is the subject in Image 2. They completely replace the two people in Video 1: face, hair or fur, build and clothes all come from the photos, nothing from the people in Video 1. Neither performer wears gloves (unless already wearing them in the photo): people show their own hands, animals keep their own paws. Ignore the photo backgrounds.

The left performer raps the vocals of Audio 1 into the microphone, lips precisely in sync with Audio 1. The right performer follows the moves of the person on the right in Video 1: reacting and gesturing on the beat, mouthing only short ad-libs.

Orange fills the whole frame, with no black areas or dark corners. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.

The AI rap variant on Seedance

On Seedance the audio is always an AI rap that the model composes itself. On the orange Hotel Lobby set it still listens to the Hotel Lobby track while composing: the reference clip is sent silent, the track is attached separately as reference audio 1, and the prompt asks for a new beat that follows its rhythm, with new words and without its melody or voices. Here the attachments are named in plain words (“reference image 1”). The example below was built for the topic shown in its label.

Hotel Lobby AI rap prompt (Seedance, 9:16, 12 s)
Use reference video 1 for the whole performance: copy its framing, timing and every movement, gesture, head turn, facial expression and mouth movement, the plain orange background and the single silver microphone hanging between the performers. Keep the camera framing of reference video 1: no zoom in, no zoom out, no pan. Do not add any movement or scene element that is not in reference video 1.

Performer on the left: the person in reference image 1. Performer on the right: the person in reference image 2. They replace the two people in reference video 1 completely: face, hair, body and clothes come from the photos, nothing from the people in reference video 1. If a photo shows none of a person's clothes (a face-only crop or bare shoulders), dress that person in plain everyday clothes that suit them, such as a simple T-shirt and jeans; never the costume, cape, coat, belts, gloves or blindfold of the people in reference video 1. Animals stay as they are, with no clothes added. Ignore the photo backgrounds.

Soundtrack: an original dark Atlanta trap song made for this video that follows reference audio 1 for its beat and rhythm: the same tempo, groove, drum pattern and bounce, timed to the moves in reference video 1. Make a new beat in that style with new lyrics; do not reuse the lyrics, melody or voices of reference audio 1. Sparse but hard-hitting: a thumping, gritty 808 that slides between notes, a piercing snare-clap with sharp hi-hats, and one eerie, Eastern-tinged melody loop. Dark, menacing, confident mood. The rap is a fast, bouncy triplet flow riding on top of the slow beat. Mix: vocals clear on top, the beat right behind them at nearly the same level, loud and punchy, never quiet and never dropping out between lines. The left performer raps an original, clean verse into the hanging microphone about: "Maya’s 30th birthday and her terrible parallel parking". Rap in the language of that topic. Clear rap vocals, lips in sync with every word, natural head nods and expressive mouth movement. The right performer is the hype partner, pointing at the rapper, reacting, smiling and adding short ad-libs on the beat. Do not use any existing song, existing lyrics or the voice of any real artist.

9:16. The orange fills the whole frame edge to edge, with no black or dark areas. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.
For an AI rap, give a name and one specific detail. This is the topic field, not the full video prompt.
For an AI rap, give a name and one specific detail. This is the topic field, not the full video prompt.Actual interface, English · 2026-10-02

Five Hotel Lobby AI prompts for two photos and no clip

Lots of video tools accept reference photos but not a reference clip. These five prompts spell out the set and the performance in words, so they paste straight into an image-to-video tool that takes two photos. They follow the rules that make the format read at a glance: one photo per side, one mic hanging from above, one voice on the verse while the other reacts, and a camera that stays put. We wrote them for this page and have not tested them, and every model reads a prompt its own way, so plan on a retry or two.

Prompt for a couple

Couple (9:16, 12 s)
Make a vertical 9:16 clip, 12 seconds long, using two reference photos of a couple. Place the subject of photo 1 on the left side of the frame and the subject of photo 2 on the right side, head to toe, and keep every face, haircut, outfit and height true to its photo. The room is a bare studio painted orange on the wall and floor, lit softly and evenly. A single silver studio mic dangles on a thin cord from above, halfway between them. The partner on the left delivers the verse into that mic with sure, easy hand moves; the partner on the right grins, points at them, sways to the beat and leans toward the mic for the closing line. The camera sits at chest height and never moves: one unbroken shot with no cuts, zoom or pan. Nobody else in the room, no on-screen words, no logos.

Prompt for two best friends who swap the mic

Best friends (9:16, 12 s)
Make a vertical 9:16 clip, 12 seconds long, from two reference photos of two grown-up best friends. Photo 1 goes on the left, photo 2 on the right, both seen from head to feet, with faces, hair, clothing and build copied exactly. The set is an orange studio with a seamless orange wall and floor, and one silver mic hanging from the ceiling on a cable between the two friends. In the first six seconds the left friend raps into the mic while the right friend nods, laughs and hypes with big arm moves; in the last six seconds the right friend takes the mic and the left friend reacts. Each friend moves in their own way, never as mirror images. Locked-off camera, a single continuous take, gentle even lighting, both pairs of feet in frame. No other people, props, captions or logos.

Prompt for a grandparent and a grown-up grandchild

Grandparent and adult grandchild (9:16, 12 s)
Make a vertical 9:16 clip, 12 seconds long, from two reference photos: photo 1 shows a grandparent, who stands on the left, and photo 2 shows their grown-up grandchild, who stands on the right. Show both from head to toe and keep each face, hairstyle, outfit and height as photographed. They share a rap duet in an empty orange studio, one silver mic hanging from the ceiling between them. The grandparent raps into the mic with calm, steady, self-assured gestures; the grandchild looks delighted, laughs, dances in time and joins the mic for the final line. Keep every movement relaxed and natural for both ages. A single fixed shot under soft studio light, with no cuts, zoom or pan, nobody else in the scene and no captions.

Prompt for drawn or anime characters

Stick to characters you created or have rights to; famous characters are owned by someone else.

Illustrated characters (9:16, 12 s)
Make a vertical 9:16 clip, 12 seconds long, from two reference images of original drawn characters. The character from image 1 stands on the left and the character from image 2 on the right, shown in full, each keeping its exact art style, palette, proportions and costume, with both drawn in that same style throughout. They perform a rap duet in a plain orange studio, a single silver mic hanging from the ceiling between them. The left character raps into the mic with lively hand gestures; the right character reacts, points and bounces on the beat. One locked camera, one continuous take, no cuts, zoom or pan. No additional characters, no lettering, no logos.

Hotel Lobby AI prompt for a dog, a cat or both

Pets were among the first Hotel Lobby clips to go viral, made with other tools (examples, credited). Give each animal one sharp photo with its whole body in frame.

Pets (9:16, 12 s)
Make a vertical 9:16 clip, 12 seconds long, from two reference photos of pets. The animal from photo 1 stands up on its back legs on the left and the animal from photo 2 stands on the right, each keeping its breed, face, eye color, coat markings, size and collar. The set is an empty orange studio with one silver mic hanging from the ceiling between them. The left animal opens and closes its mouth on the beat and waves its front paws as though rapping into the mic; the right animal bobs its head, glances at the camera and dances in time. Keep the anatomy real: four legs and one tail per animal, both visible from the tips of the ears to the paws. A single fixed, unbroken shot. No humans, no extra animals, no text.

Our studio needs none of these, for people or for pets, because the prompt is already built in; the pet rap video page opens it set up for a pet and its partner. We have tested a dog beside a person. Cats, and two pets together, are still untested.

Why each instruction is in the prompt

  1. Performance. The opening paragraph gives Video 1 control of framing, timing and every movement: gestures, head turns, faces, mouth shapes, plus the plain orange backdrop and the one silver mic hanging between the two. It freezes the camera (no push-in, pull-back or pan) and forbids anything the clip does not contain.
  2. Audio. Audio 1 is the entire soundtrack and is left untouched: no rewritten lyrics, no re-sung lines, no added music or voices.
  3. Who is who. Each image is bound to one side, and the two subjects fully take the place of the two performers in the clip: face, hair or fur, body and clothes all come from the images, and their backgrounds are dropped. The wording is “the subject in Image 1” rather than man or woman, so it suits a person or a pet and never argues with the photo. It also bans gloves, because the two dancers in our reference clip wear them and the model used to paint them onto hands and paws.
  4. Parts. The left performer raps the vocals of Audio 1 with lips in sync; the right performer copies the sidekick in the clip, reacting and gesturing on the beat with only brief ad-libs.
  5. Safety rails. Orange all the way to every edge with no dark corners, no split screen, no extra people, no captions, and the same faces from first frame to last. These lines prevent the failures we saw most.

Porting the prompt to another video model

  • You need a model that accepts reference images, a reference video and a reference audio track, and that keeps the track. Wan 3.0 does. ByteDance Seedance 2.x accepts all three, but in our tests it composed new music.
  • Refer to attachments the way your tool does. We number them in upload order (Image 1, Image 2, Video 1, Audio 1); other tools want @Image1 or their own tags. Keep the order fixed: left photo first, right photo second.
  • Leave the performers unnamed. “The subject in Image 1” fits whatever is in the picture, human or animal. Our studio only changes the photo layout (two photos or one of both) and, when you upload Your clip, the moves.
  • The frame shape is not in the text; choose 9:16, 16:9 or 1:1 in the tool’s settings.
  • No reference clip? Describe the set and the performance in words instead of pointing to Video 1, as in the five prompts above. Dreamina and similar tools publish their own text-only versions. We have tested only the two prompts at the top of this page.

Fix-it table: what goes wrong and why

You seeProbable reasonTry this
A face turns into a strangerThe photo shows a side view, sunglasses or a tiny faceSwap in a front-on, bright photo framed from the waist up
Left and right trade placesNothing in the prompt binds each photo to a sideName the left and right performer with their image
An extra person walks inPeople from the clip or a photo background get copiedKeep “no extra people” and tell it to ignore photo backgrounds
Text pops up on screenThe model invents captionsKeep the line that bans captions and on-screen text
Mouths drift off the audioNo audio attached, or the track opens on silenceAttach the track and start it on a word
Black borders around the orangeThe prompt asks for a camera move the clip never makesKeep the clip’s framing and ask for orange edge to edge
Different music replaces the trackThe model writes its own soundtrack (Seedance 2.x did for us)Pick a model that keeps the reference track, such as Wan 3.0

Or let the studio write it

In Hotel Lobby Video the prompt, the reference clip and the reference track are already wired in. You add a pair of photos and press create. A 12-second take on Wan 3.0 at 480p costs 6 credits.

Create your Hotel Lobby video

FAQ

What makes a good Hotel Lobby AI prompt?

It fixes the orange set, the single hanging mic, which photo goes on which side and what each person does. The Wan 3.0 prompt above is the one our studio sends for every take.

Which models does the prompt work with?

We run the first with Wan 3.0, which keeps the reference track, and the AI rap version with Seedance 2.x, which writes its own music. Both accept reference images, a reference video and reference audio. Other models need their own tagging and may not take all three.

Can I paste the couple prompt into Dreamina or Kling with two photos from my camera roll?

That is what the five photo-only prompts are for: image-to-video tools that take two reference photos. Rename the photos the way your tool expects and keep left and right fixed. We have not tested them outside our own studio.

I want my dog and me in it. Do I need the pet prompt?

Not on Hotel Lobby Video: the [pet rap video page](/dog-and-cat-rap-video) is set up for a pet and its partner, with no prompt to write. A dog beside a person is tested; cats and two pets together are not yet.

Why does my prompt keep producing new music instead of the Hotel Lobby track?

Some models compose their own soundtrack whatever you ask; Seedance 2.x did in our tests. Use a model that keeps a reference track, such as Wan 3.0, and attach the track as audio.

Do I have to write any prompt at all?

No. Hotel Lobby Video builds the prompt into the studio, so you only add the photos.

References

Hotel Lobby Video is an independent app with no ties to Quavo, Takeoff, Migos, Quality Control Music, Motown or COLORS.