CIRCE: Building a Complete AI Music Artist from Zero

Circe music video screenshot, she is standing in the dark candles hall

The question

Every musician, every label, every artist account is running the same race: face, music, video, presence. The question I wanted to answer: what happens when none of those require a real person?

CIRCE is a personal project built to test that question directly. The brief I wrote for myself: build a complete solo music artist from nothing (face, character, original music, full music video, live platforms) using only AI production tools under human creative direction. No studio. No musicians. No real person at any stage.

The result is live on YouTube. This is how it was made.

Six photorealistic film stills arranged in a 3×2 grid. Location: a dark editing workstation setup. Large monitor displaying the Magnific AI interface — music video footage being processed, VFX layers visible on screen. Second screen showing raw footage reference. Keyboard, stylus, headphones on the desk. The room is dark except for screen glow — blue-white light on both subjects. Two subjects: the platinum-blonde woman seated at the workstation, in dark comfortable clothing — dark oversized jacket or sweater, gothic jewelry still visible at the neck. She faces the BTS camera, animated. Beside or slightly behind her: a second person (male) — input image defines his appearance — watching the screen or glancing at the BTS camera. Photo 1 (medium): both subjects at the workstation. Magnific interface visible on the screen behind them — footage processing. She faces the BTS camera with energy. He watches the screen. Photo 2 (medium): she speaks, animated — genuine excitement. Screen glow on both of them. His face visible in profile or three-quarter. Photo 3 (medium close-up): she faces BTS camera. Screen in background out of focus — VFX processing layers visible as soft colour shapes. Photo 4 (medium): both visible. She continues speaking — mid-line. His expression: amused, tired in a satisfied way. Screen glow. Photo 5 (close-up): her face. Screen light from below and one side. Animated — the most expressive she has been in any BTS clip. Photo 6 (medium): she finishes the line. She glances at the screen, then back at the BTS camera. He is still watching the screen. The Magnific interface shows a processed MV clip on the monitor.

What exists

In five days of production, May 15 to 19, 2026, the following was built from zero:

  • A locked character identity: five visual looks, a full character bible, face and body reference plates calibrated across more than one hundred generated images
  • A logo: granite carving aesthetic, serpent motif, CIRCE wordmark
  • An original symphonic gothic metal track: “Candle’s Hall” (operatic soprano, full orchestral arrangement, original lyrics written before a single generation prompt was submitted)
  • A 4-minute-34-second music video: atmospheric cinematic opening, full track music video with performance and narrative woven together, and a behind-the-scenes disclosure section assembled from more than thirty AI-generated clips in Premiere Pro
  • Three live platforms: YouTube at @circe.aimusic, Instagram, and TikTok
  • A social content pipeline: two BTS content entries complete, caption voice system, ongoing

The identity problem

The hardest engineering problem in AI character production is not generation quality. It is consistency. A character generated across one hundred sessions, with prompts written in different moods on different days, becomes three different people. The face drifts. The hair drifts. The energy drifts. Without a system, what you end up with is a gallery of attractive strangers.

The solution is a character bible written before the first image is generated. Not a mood board. A technical specification: hair color described with explicit exclusion terms (“warm honey-ash root fading to cold ice-white, NOT silver, NOT grey”) because “platinum blonde” and “silver” produce different outputs and the model will pick the wrong one if you leave it room. Eye color specified by contrast value, not just hue (“near-black dark brown, NOT green, NOT hazel, NOT light brown”) because close-up lighting consistently drifts eye color toward green unless you name the wrong answers explicitly. Ethnicity described through physical trait combinations rather than nationality (“sharp cheekbones from one origin, oblique almond eye shape from another, full sculpted lips, warm-toned medium fair skin, genuinely unplaceable”) because naming a country gives the model permission to collapse the whole face into a single cultural aesthetic.

The second rule: face and body lock before any wardrobe. The first generation is always a neutral calibration plate: minimal clothing, natural makeup, bright even lighting, nothing competing with the face. If the face is wrong, fix it here. Not after ten wardrobe shots. Once the neutral plate is approved, wardrobe is built as a separate prompt layer on top of the locked identity. Identity and styling are never mixed into the same description block.

The result: five complete looks, more than one hundred generated images across sessions, consistent enough to read as the same person in every frame.

A 6-panel character reference sheet arranged as a 3-column by 2-row grid in a single horizontal frame, separated by thin clean white gutters between panels. Each panel shows the same single character — a slim tall woman with lean feminine proportions: long platinum blonde hair loose and slightly wavy, falling past the shoulders and collarbone, individual strand-by-strand definition, flyaways catching the key light, the platinum reading warm-root-to-cold-ice-white blonde NOT silver NOT grey. Mixed-ambiguous multi-origin ethnicity: sharp high angular cheekbones, full sculpted lips with a pronounced cupid's bow upper lip and full rounded lower lip, near-black dark brown eyes — almost black, deeply dark, NOT green NOT hazel NOT light brown, the iris nearly indistinguishable from the pupil at normal distance — heavy-lidded, the oblique outer corner pull from the bone structure creating an inherently cat-eye shape without liner, warm-toned medium olive-fair skin — medium fair NOT deep tan NOT dark, a golden-olive undertone under a pale surface — genuinely unplaceable. Gothic dramatic makeup: full-face dark contour with deep shadow sculpted along the hollows of the cheeks and the temples, the shadow darkening the face theatrically, powder-matte skin finish reading ghost-pale under the key light — the skin matte and theatrically pale, NOT dewy, NOT glossy. Heavy liner applied thickly along the full upper and full lower lash lines creating a complete dark ring around the eye. Near-black charcoal shadow packed from the lash line across the entire lid and into the outer corner and crease. A deep vampiric dark wine-to-near-black lip — reading almost black in the deepest shadow and deep dark crimson in the light, matte finish, full coverage edge to edge, the lip fullness clearly defined at the cupid's bow and the lower lip. Wearing a structured strapless black fitted bodice in heavy black crepe, internal boning giving the bodice clean architectural structure — NOT a corset with visible external lacing, a clean smooth strapless silhouette — the hem ending at the upper thigh as a short structured fitted skirt. Draped over the shoulders: a floor-length sheer black organza cape, the fabric attached at each shoulder point and falling straight to the floor, completely open at the center front — the two panels of the cape falling away to each side and framing the body from the front like dark wings, the cape fabric sheer and translucent, the key light reading through the organza as a dark shimmer, real organza weight and drape with the hem pooling slightly at the floor. Knee-high structured black leather boots with a thick architectural gothic heel — NOT the stiletto of Look 1, a wider and taller block heel with a gothic silhouette, pointed toe, boot shaft ending exactly at the knee, the leg between the boot top and the skirt hem fully visible through the open front of the cape. At the neck: a wide dramatic gothic choker in dark oxidized silver spanning from the collarbone to the jaw — architectural in scale, with hanging drop pendants along its lower edge in articulated dark metal, the surface of the choker heavily worked dark metal. A long dark faceted stone pendant on a thin dark chain falling from the neck to the sternum, visible against the strapless bodice above the bust line. Multiple dark stone rings — black onyx or similar deep faceted stone set in dark oxidized silver settings — on both hands, two to three rings per hand. Long drop earrings in dark oxidized silver falling from the ear to the shoulder level, visible against the platinum hair at the sides of the face.

The music

The track direction came before any generation prompt. Not “make a gothic metal song.” A full creative brief: the emotional arc from near-silence to full orchestral eruption to a stripped bridge of solo cello and voice; the singer shifting from close-mic half-whisper in the verses to full operatic power at the chorus; the lyric frame addressing the listener directly, in second person, from a position of dark control.

The lyrics were written first, before the generation tools opened. The lyrics are not decoration on top of the project. They are the conceptual spine of everything that follows. The short film opening was scripted around the chorus hook. The video’s BTS disclosure scene was written to echo the lyric logic. The campaign’s bio, the channel description, the release captions: all pull from the same language.

The chorus:

You should not have looked
You cannot look away
You came of your own will
Now you’re going to stay

“You came of your own will” became the Instagram bio. It became the YouTube channel description. It became the logic of the entire AI disclosure strategy. The track and the visual identity are the same idea expressed in two different forms. That coherence does not happen by accident. It happens because the creative brief precedes every tool, not the other way around.

| Panel | Moment | |---|---| | 1 — Close face | Her face, centred. She is looking directly into the lens. No upward gaze — direct. The previous warmth is gone. She is describing something. | | 2 — Detail | Her eyes only. The iris. The candlelight in it. She has not looked away. | | 3 — Mid | Pulled back slightly — face and upper chest. She sings and the camera is completely still. She is the only thing moving. | | 4 — Low angle | Slight upward angle. Hard shadow on one side. She is above the camera. The words come from above. | | 5 — Detail | Her lips again — different word, different shape. "Watching." The W forming. | | 6 — Dramatic | Full face. Both eyes in frame. Hard shadow line. She holds the lens as the verse ends. She does not release the gaze before the cut. |

The video

The video opens without music. An empty palatial hall: chessboard marble floor, wall sconces already lit, no one there. A cut to a misty forest clearing at night. Then she appears, and from that point the edit does not pause. The performance sections, the narrative beats, and the abstract B-roll are all woven together rather than separated into labeled chapters. The structure that was planned on paper ended up tighter and more integrated in the edit.

Two tools, two different jobs. The first handles cinematic clips with no audio dependency: environment plates, abstract montage, the opening sequence. These clips need high visual fidelity and precise camera movement, no audio input, longer generation cycles, less compression. The second handles audio-driven clips: every singing section, every BTS moment, every clip where the character’s physical delivery needs to track the vocal. Each of these receives its own trimmed audio segment as input, exported directly from the Premiere timeline before generation. The motion follows the sound.

Every generated clip is fifteen seconds or shorter, regardless of which tool produced it. The video is built from more than thirty individual clips. This is not a constraint to work around. It is the correct production unit. Shorter clips generate more cleanly, and the fifteen-second limit forces the edit structure to be explicit from the start. You cannot hide behind a long take. Every cut is a decision made before generation, not after.

Several visual moments in the video are worth naming specifically because they were not predetermined. They came out of the generation process and were kept as improvements on the original plan. The chorus runs a triple-image effect: three instances of the same character, side by side, synchronised. It was generated as a single clip and landed in one pass. The floor crack sequence (her hands pressed flat to the marble as it splits open beneath her) is more visceral than any equivalent described in the brief. The mirror does not hold her reflection a beat too long; it shatters completely. She holds the serpent directly in her hands rather than having it coiled in the background. Each of these is a moment where the generation went further than specified and the edit kept it.

The voice architecture runs the same split logic. The singing voice and the speaking voice are generated separately with separate profiles. The singing voice (operatic soprano, close-mic in the verses, full power in the chorus) is produced as part of the music generation. The speaking voice used in the behind-the-scenes section is warmer, more direct, a different register entirely, generated with its own configuration. The contrast between the two is deliberate. Performance is cold and operatic. BTS is warm and present. That difference makes both feel more real.

BTS credibility depends on production detail. The behind-the-scenes section was prompted with the same specificity as the cinematic sections: crew members visible in the forest, a monitor and DIT station in the background, LED work lights alongside the candles in the palatial interior. Without those details, BTS reads as staged. With them, the production reads as real, which is the condition under which the disclosure that follows can do what it needs to do.

Six photorealistic film stills arranged in a 3×2 grid. Location: forest clearing at night. Production set — fog machines off, working light only. Flat ambient light. Film crew and equipment visible: cinema camera on tripod, LED panel on stand, cables on ground, monitor station, crew members in dark puffer jackets. Subject: platinum-blonde woman in a dark oversized puffer jacket worn over a black dress. Gothic jewelry at the neck. Photo 1 (wide): full production set visible. Subject in foreground facing left. Crew and camera rig in mid-ground. Trees behind. Photo 2 (medium): subject facing the camera. Crew member and monitor station out of focus behind her. Photo 3 (medium close-up): subject facing camera, relaxed expression. Equipment blurred in background. Photo 4 (medium close-up): subject speaking, natural expression. Crew movement blurred behind her. Photo 5 (close-up): face, slight upward curve at the corner of her mouth. Edge of a lighting rig visible at frame left. Photo 6 (medium): subject looking at camera. Crew member carrying equipment visible walking behind her.

The reveal

The standard approach to AI disclosure is a disclaimer. A footnote. A press release after launch explaining the tool list.

That approach treats disclosure as a liability to manage. CIRCE treats it as a creative decision, and builds it in three beats.

The first beat is a throwaway line in the forest BTS. She is standing in her puffer jacket between takes, crew and monitors visible behind her, and she says: “People always ask if it’s cold. It’s always cold.” Warm, dry, present. The register has completely shifted from the performance footage that preceded it. That shift is the setup.

The second beat is in the palatial interior. A makeup artist finishes a touch-up. She turns to the BTS camera:

“Worth it. I’m not real, but neither is whatever you felt watching this.”

The makeup touch-up is the credibility anchor for this line. The viewer has watched four minutes of what reads as a real production: crew, gear, a palatial location, a forest shoot, a musician between takes. The makeup artist interaction is entirely human, grounded, specific. It is what makes the disclosure land as something other than a mechanical statement. She speaks after a human moment, not instead of one. And “worth it,” the first word, does more work than the whole sentence that follows. It makes the line a judgment, not a confession.

The third beat is the one that changes the frame entirely. The next cut is to a dark editing suite: she is sitting at a workstation, Premiere Pro open on the monitor, footage from the video visible on screen. A man is standing behind her. She turns to the camera and delivers the final lines of the video:

“As he has worked on this for 16 hours straight, I took the mouse out of his hand. We’re almost done.”

The disclosure is now not just an emotional statement. It is a collaborator reveal. She claims authorship of the project. The person who actually built it is visible in the frame behind her, named only by his hours. The warmth and dry humor of the line, the complete tonal opposite of the cold operatic character in the first four minutes, is the final proof that two registers were built and held throughout. The video ends on the stone-carved logo, the serpent coiled at its base, and cuts to black.

Stage 6 of the campaign (the disclosure) is not the end of the project. It is the beginning of the conversation this project was always designed to start.

Six photorealistic film stills arranged in a 3×2 grid. Location: grand palatial stone interior. Stone wall close — the ancient worn markings visible on the surface. Low candlelight. Subject: same platinum-blonde woman. She stands at the stone wall, her palm held flat near the surface — not touching the stone, a few centimetres from it. The markings near her hand shift in response to her proximity. She watches them shift. She does not trace them. She does not need to touch. Photo 1 (medium): she stands at the stone wall, palm raised near the surface but not touching. Her face in three-quarter profile. The markings on the wall are worn and still. Photo 2 (close-up): her palm near the stone — close but not touching. Her dark onyx rings visible. The markings adjacent to her hand begin to catch light differently. Photo 3 (close-up): the marking nearest her palm — slightly rotated. A faint groove in the stone dust where it moved. Her hand has not moved. Photo 4 (medium close-up): her face watching the wall. Even. She is reading what she sees, not surprised by it. Photo 5 (close-up): two markings now catching light — both near her hand. The stone unchanged except at those points. Photo 6 (medium): she lowers her hand slowly. The markings hold their new positions. She turns slightly away from the wall.

Three engineering problems

Safety filter vocabulary. Symphonic gothic metal sits close to territory the generation safety systems watch: expressive wardrobe, intense gaze descriptions, atmospheric settings that read as intimate. The first version of the Look 2 character sheet was flagged before a single image generated. The fix was not to work around the filter but to replace the vocabulary entirely. Every term that could read as sexualizing context was mapped to a fashion and editorial equivalent: “sultry” became “dramatic,” descriptors for bare skin were removed, gaze descriptions shifted to editorial register. The substitution is not cosmetic. Fashion and editorial vocabulary signals the correct generation context to the model. It sits in a different category than the content the safety system is designed to catch. Writing in that register produces better outputs, not just safer ones. The vocabulary list built during this project now travels into every character generation job.

Phone quality in BTS video generation. The social content pipeline uses a selfie and BTS aesthetic: phone-quality footage, natural and unpolished. The problem: if the reference image used to generate a BTS reel is photorealistic editorial quality, the video generation model anchors on that quality and ignores the phone descriptor in the video prompt. Text and image reference have to agree. The fix is to bake phone sensor quality directly into the reference image generation: digital grain, compressed dynamic range, sensor bloom near light sources, imperfect auto-exposure, cool phone color rendering. The moodboard grid has to look like a phone shot before the reel prompt asks for one. If the reference looks editorial, the output will too, regardless of what the text says.

Motion specification in AI video. Still or near-still characters in video generation produce slideshow quality. The model interpolates nothing and defaults to a subtle zoom or grain pass. Any clip where the character appears must specify explicit named motion: a slow head turn, a deliberate step forward, a cape shifting as weight transfers. The motion description has to be physical and measurable, not emotional. “She stares intensely” produces stillness. “Her head turns three degrees left over four seconds, eyes holding the lens” produces motion. The constraint is consistent across every video generation tool used in this project and applies equally to cinematic and audio-driven pipelines.

Screenshot of magnific.com workflow node system in Spaces

What this project proves

A complete artist launch (character, music, video, platforms) produced in five days by one person with no collaborators, no studio, and no budget. The traditional equivalent requires a character designer, a photographer, a composer, a recording studio session, a music video director, a camera operator, a location scout, a video editor, and a social media manager. CIRCE required none of those. It required a creative brief, a production methodology, and the discipline to enforce both across every tool and every decision.

The quality gap between a generic AI music video and this project is not a tool gap. Every tool used here is available to anyone with a subscription. The gap is entirely in the creative direction: the character bible, the lyric concept, the dual-pipeline video structure, the two-register voice system, the BTS credibility techniques, the disclosure built into the film. None of those are AI outputs. They are decisions made before any generation begins.

CIRCE is not finished. The release campaign is still running. A second track is planned. The social content pipeline is active. What the five-day production window proves is the starting point: one brief, one methodology, zero collaborators, a complete artist.

The production method applies to any project that requires a character, a voice, a video, or an identity built from nothing.

Six photorealistic film stills arranged in a 3×2 grid. Location: grand palatial stone interior. Cold stone floor. Candlelight from one side. Subject: platinum-blonde woman in black structured strapless mini dress, floor-length sheer black organza cape, gothic choker with drop pendants, dark onyx rings. Heavy gothic makeup. She holds one arm extended — palm up. A dark iridescent serpent, green-black scales, coils slowly up her extended forearm. She is watching it. She is not afraid. She is not performing. She brought it here. Photo 1 (medium): she stands in the interior, one arm extended at her side, palm up. The serpent is at her wrist — beginning to coil. Her eyes are on it. Photo 2 (medium): the serpent has coiled further up her forearm. She watches, arm still. Candlelight on both of them. Photo 3 (close-up): her forearm and the serpent — iridescent scales against her skin. Her hand open, relaxed. Not gripping. Photo 4 (medium close-up): her face in three-quarter profile as she watches the serpent. No alarm. No expression of effort. She is simply present with it. Photo 5 (close-up): the serpent's body coiled at her elbow. Her dark onyx rings visible. Scale texture against her skin. Photo 6 (medium): wide enough to see her full upper body. The serpent fully coiled to her elbow. She has not moved. She is still watching. The candlelight catches the iridescent surface of the scales.

Let’s Create Together!

Loved this project? Got questions or want to collaborate on something amazing? I’d love to hear from you!
just drop a quick message — let’s start a conversation!