CIRCE: Building a Complete AI Music Artist from Zero
The question
Every musician, every label, every artist account is running the same race: face, music, video, presence. The question I wanted to answer: what happens when none of those require a real person?
CIRCE is a personal project built to test that question directly. The brief I wrote for myself: build a complete solo music artist from nothing (face, character, original music, full music video, live platforms) using only AI production tools under human creative direction. No studio. No musicians. No real person at any stage.
The result is live on YouTube. This is how it was made.
What exists
In five days of production, May 15 to 19, 2026, the following was built from zero:
- A locked character identity: five visual looks, a full character bible, face and body reference plates calibrated across more than one hundred generated images
- A logo: granite carving aesthetic, serpent motif, CIRCE wordmark
- An original symphonic gothic metal track: “Candle’s Hall” (operatic soprano, full orchestral arrangement, original lyrics written before a single generation prompt was submitted)
- A 4-minute-34-second music video: atmospheric cinematic opening, full track music video with performance and narrative woven together, and a behind-the-scenes disclosure section assembled from more than thirty AI-generated clips in Premiere Pro
- Three live platforms: YouTube at @circe.aimusic, Instagram, and TikTok
- A social content pipeline: two BTS content entries complete, caption voice system, ongoing
The identity problem
The hardest engineering problem in AI character production is not generation quality. It is consistency. A character generated across one hundred sessions, with prompts written in different moods on different days, becomes three different people. The face drifts. The hair drifts. The energy drifts. Without a system, what you end up with is a gallery of attractive strangers.
The solution is a character bible written before the first image is generated. Not a mood board. A technical specification: hair color described with explicit exclusion terms (“warm honey-ash root fading to cold ice-white, NOT silver, NOT grey”) because “platinum blonde” and “silver” produce different outputs and the model will pick the wrong one if you leave it room. Eye color specified by contrast value, not just hue (“near-black dark brown, NOT green, NOT hazel, NOT light brown”) because close-up lighting consistently drifts eye color toward green unless you name the wrong answers explicitly. Ethnicity described through physical trait combinations rather than nationality (“sharp cheekbones from one origin, oblique almond eye shape from another, full sculpted lips, warm-toned medium fair skin, genuinely unplaceable”) because naming a country gives the model permission to collapse the whole face into a single cultural aesthetic.
The second rule: face and body lock before any wardrobe. The first generation is always a neutral calibration plate: minimal clothing, natural makeup, bright even lighting, nothing competing with the face. If the face is wrong, fix it here. Not after ten wardrobe shots. Once the neutral plate is approved, wardrobe is built as a separate prompt layer on top of the locked identity. Identity and styling are never mixed into the same description block.
The result: five complete looks, more than one hundred generated images across sessions, consistent enough to read as the same person in every frame.
The music
The track direction came before any generation prompt. Not “make a gothic metal song.” A full creative brief: the emotional arc from near-silence to full orchestral eruption to a stripped bridge of solo cello and voice; the singer shifting from close-mic half-whisper in the verses to full operatic power at the chorus; the lyric frame addressing the listener directly, in second person, from a position of dark control.
The lyrics were written first, before the generation tools opened. The lyrics are not decoration on top of the project. They are the conceptual spine of everything that follows. The short film opening was scripted around the chorus hook. The video’s BTS disclosure scene was written to echo the lyric logic. The campaign’s bio, the channel description, the release captions: all pull from the same language.
The chorus:
You should not have looked
You cannot look away
You came of your own will
Now you’re going to stay
“You came of your own will” became the Instagram bio. It became the YouTube channel description. It became the logic of the entire AI disclosure strategy. The track and the visual identity are the same idea expressed in two different forms. That coherence does not happen by accident. It happens because the creative brief precedes every tool, not the other way around.
The video
The video opens without music. An empty palatial hall: chessboard marble floor, wall sconces already lit, no one there. A cut to a misty forest clearing at night. Then she appears, and from that point the edit does not pause. The performance sections, the narrative beats, and the abstract B-roll are all woven together rather than separated into labeled chapters. The structure that was planned on paper ended up tighter and more integrated in the edit.
Two tools, two different jobs. The first handles cinematic clips with no audio dependency: environment plates, abstract montage, the opening sequence. These clips need high visual fidelity and precise camera movement, no audio input, longer generation cycles, less compression. The second handles audio-driven clips: every singing section, every BTS moment, every clip where the character’s physical delivery needs to track the vocal. Each of these receives its own trimmed audio segment as input, exported directly from the Premiere timeline before generation. The motion follows the sound.
Every generated clip is fifteen seconds or shorter, regardless of which tool produced it. The video is built from more than thirty individual clips. This is not a constraint to work around. It is the correct production unit. Shorter clips generate more cleanly, and the fifteen-second limit forces the edit structure to be explicit from the start. You cannot hide behind a long take. Every cut is a decision made before generation, not after.
Several visual moments in the video are worth naming specifically because they were not predetermined. They came out of the generation process and were kept as improvements on the original plan. The chorus runs a triple-image effect: three instances of the same character, side by side, synchronised. It was generated as a single clip and landed in one pass. The floor crack sequence (her hands pressed flat to the marble as it splits open beneath her) is more visceral than any equivalent described in the brief. The mirror does not hold her reflection a beat too long; it shatters completely. She holds the serpent directly in her hands rather than having it coiled in the background. Each of these is a moment where the generation went further than specified and the edit kept it.
The voice architecture runs the same split logic. The singing voice and the speaking voice are generated separately with separate profiles. The singing voice (operatic soprano, close-mic in the verses, full power in the chorus) is produced as part of the music generation. The speaking voice used in the behind-the-scenes section is warmer, more direct, a different register entirely, generated with its own configuration. The contrast between the two is deliberate. Performance is cold and operatic. BTS is warm and present. That difference makes both feel more real.
BTS credibility depends on production detail. The behind-the-scenes section was prompted with the same specificity as the cinematic sections: crew members visible in the forest, a monitor and DIT station in the background, LED work lights alongside the candles in the palatial interior. Without those details, BTS reads as staged. With them, the production reads as real, which is the condition under which the disclosure that follows can do what it needs to do.
The reveal
The standard approach to AI disclosure is a disclaimer. A footnote. A press release after launch explaining the tool list.
That approach treats disclosure as a liability to manage. CIRCE treats it as a creative decision, and builds it in three beats.
The first beat is a throwaway line in the forest BTS. She is standing in her puffer jacket between takes, crew and monitors visible behind her, and she says: “People always ask if it’s cold. It’s always cold.” Warm, dry, present. The register has completely shifted from the performance footage that preceded it. That shift is the setup.
The second beat is in the palatial interior. A makeup artist finishes a touch-up. She turns to the BTS camera:
“Worth it. I’m not real, but neither is whatever you felt watching this.”
The makeup touch-up is the credibility anchor for this line. The viewer has watched four minutes of what reads as a real production: crew, gear, a palatial location, a forest shoot, a musician between takes. The makeup artist interaction is entirely human, grounded, specific. It is what makes the disclosure land as something other than a mechanical statement. She speaks after a human moment, not instead of one. And “worth it,” the first word, does more work than the whole sentence that follows. It makes the line a judgment, not a confession.
The third beat is the one that changes the frame entirely. The next cut is to a dark editing suite: she is sitting at a workstation, Premiere Pro open on the monitor, footage from the video visible on screen. A man is standing behind her. She turns to the camera and delivers the final lines of the video:
“As he has worked on this for 16 hours straight, I took the mouse out of his hand. We’re almost done.”
The disclosure is now not just an emotional statement. It is a collaborator reveal. She claims authorship of the project. The person who actually built it is visible in the frame behind her, named only by his hours. The warmth and dry humor of the line, the complete tonal opposite of the cold operatic character in the first four minutes, is the final proof that two registers were built and held throughout. The video ends on the stone-carved logo, the serpent coiled at its base, and cuts to black.
Stage 6 of the campaign (the disclosure) is not the end of the project. It is the beginning of the conversation this project was always designed to start.
Three engineering problems
Safety filter vocabulary. Symphonic gothic metal sits close to territory the generation safety systems watch: expressive wardrobe, intense gaze descriptions, atmospheric settings that read as intimate. The first version of the Look 2 character sheet was flagged before a single image generated. The fix was not to work around the filter but to replace the vocabulary entirely. Every term that could read as sexualizing context was mapped to a fashion and editorial equivalent: “sultry” became “dramatic,” descriptors for bare skin were removed, gaze descriptions shifted to editorial register. The substitution is not cosmetic. Fashion and editorial vocabulary signals the correct generation context to the model. It sits in a different category than the content the safety system is designed to catch. Writing in that register produces better outputs, not just safer ones. The vocabulary list built during this project now travels into every character generation job.
Phone quality in BTS video generation. The social content pipeline uses a selfie and BTS aesthetic: phone-quality footage, natural and unpolished. The problem: if the reference image used to generate a BTS reel is photorealistic editorial quality, the video generation model anchors on that quality and ignores the phone descriptor in the video prompt. Text and image reference have to agree. The fix is to bake phone sensor quality directly into the reference image generation: digital grain, compressed dynamic range, sensor bloom near light sources, imperfect auto-exposure, cool phone color rendering. The moodboard grid has to look like a phone shot before the reel prompt asks for one. If the reference looks editorial, the output will too, regardless of what the text says.
Motion specification in AI video. Still or near-still characters in video generation produce slideshow quality. The model interpolates nothing and defaults to a subtle zoom or grain pass. Any clip where the character appears must specify explicit named motion: a slow head turn, a deliberate step forward, a cape shifting as weight transfers. The motion description has to be physical and measurable, not emotional. “She stares intensely” produces stillness. “Her head turns three degrees left over four seconds, eyes holding the lens” produces motion. The constraint is consistent across every video generation tool used in this project and applies equally to cinematic and audio-driven pipelines.
What this project proves
A complete artist launch (character, music, video, platforms) produced in five days by one person with no collaborators, no studio, and no budget. The traditional equivalent requires a character designer, a photographer, a composer, a recording studio session, a music video director, a camera operator, a location scout, a video editor, and a social media manager. CIRCE required none of those. It required a creative brief, a production methodology, and the discipline to enforce both across every tool and every decision.
The quality gap between a generic AI music video and this project is not a tool gap. Every tool used here is available to anyone with a subscription. The gap is entirely in the creative direction: the character bible, the lyric concept, the dual-pipeline video structure, the two-register voice system, the BTS credibility techniques, the disclosure built into the film. None of those are AI outputs. They are decisions made before any generation begins.
CIRCE is not finished. The release campaign is still running. A second track is planned. The social content pipeline is active. What the five-day production window proves is the starting point: one brief, one methodology, zero collaborators, a complete artist.
The production method applies to any project that requires a character, a voice, a video, or an identity built from nothing.