"Not a song. Not a chapter.
It's the same moment." from the desk, Fractured Vows Saga
Story and Song, Written Together
Most soundtracks are created after the book is finished โ a composer listens to the story and writes music to match it. My process is completely different.
The songs and the story are born at the same time.
While I am deep in writing a chapter โ building a scene, shaping a character's emotion, finding the exact moment where something breaks or ignites โ lines start coming. Not prose lines. Song lines. A phrase with rhythm. A feeling that is too compressed for a paragraph โ but too big for a sentence. I write them down on the side, separate from the manuscript, and keep writing the chapter.
Over time those fragments accumulate. A chorus forms from something a character almost said out loud but didn't. A verse comes from the internal monologue I chose not to put in the book. A bridge appears from the space between two scenes โ the silence, the tension, the thing neither character will name.
By the time a song is finished, it has gone through the same process as a chapter โ drafted, revised, shaped, and reshaped until it earns its place in the story.
The book and the soundtrack are not two separate projects. They are one work, told in two languages โ prose and song.
What Lyric Writing Actually Requires
Most people misunderstand what songwriting is. A lyric is not a diary entry. It is not a character's dialogue written down. It is a completely different form โ and it has to do many things at once that prose never has to do.
Every line I write must:
- Scan โ the syllables have to land on the beat naturally, not forced. A line that reads beautifully on paper can collapse completely when sung.
- Breathe โ a singer needs space, phrasing, places to inhale. A lyric that gives no room to breathe is unsingable, no matter how good the words are.
- Rhyme โ not artificially, but in a way that feels inevitable. A forced rhyme breaks the spell. The right rhyme makes the listener feel the line was always going to end exactly there.
- Compress โ an emotion that takes three paragraphs of prose has to live in four lines. Every word earns its place or it does not come.
- Repeat โ a chorus has to be worth hearing six times without feeling stale. The words have to deepen on each repetition, not flatten.
- Land โ the last line of every section has to hit. Not just conclude. Hit.
I rewrite every lyric many times. Words that work on paper often collapse when sung. Rhymes that feel clever in silence feel forced when they are carried by melody. The adjustment is constant โ finding the word that means the right thing, and sits right in the mouth, and lands on the right beat, and carries the right emotional weight.
The Specific Struggle of Hindi Lyrics
Writing song lyrics in Hindi adds an entirely different layer of difficulty โ and the Hindi songs in this collection are not translations. They are original compositions written from scratch in Hindi, which means every challenge of lyric writing happens in a second language, inside a completely different phonetic system.
In Hindi, the sound of a word in the mouth matters as much as its meaning on the page. The same meaning can be expressed multiple ways โ but only one of those ways will sit correctly on a beat, breathe correctly in a phrase, and land correctly on the ear.
Small things become enormous decisions. Whether a word ends in a hard "d" sound or a softer "dh" sound changes how the line lands completely. Whether a vowel is held long or cut short changes the entire feel of the phrase. A word that looks right written down can sound wrong when sung โ too soft where the melody needs weight, too sharp where the melody needs warmth.
I work through multiple attempts of the same line โ sometimes four or five versions of a single phrase โ testing each one against the melody until the sound, the meaning, and the rhythm all align at once.
Each Hindi song was regenerated between 20 and 30 times before it was right. Not tweaked โ regenerated from the beginning. Because one wrong syllable, one sound that sits incorrectly on a beat, one vowel that doesn't breathe the way the phrase needs it to โ and the whole song loses its truth. I would rather start again than release something that is almost right. Almost right is wrong.
This is not something AI does. This is something I do, one word at a time, one attempt at a time, one regeneration at a time, until the line is right.
Vocal Direction Every Breath Intentional
Once the lyrics are written, the work of directing the performance begins. And this is where most people โ even those who understand songwriting โ underestimate the depth of what is required.
For every song I produce, I write complete vocal direction notes โ not just what to sing, but how to sing it. Where the voice drops. Where it cracks. Where it holds. Where the breath comes before the word and where it comes after. Where anger lives underneath the surface and never breaks through. Where restraint is the entire performance.
I mark every lyric sheet with breath positions, tonal direction, and emotional instruction. Things like:
- "Raspy here โ light grit, not full roughness." โ The difference between these two is a complete change in how the line feels to the listener.
- "Anger underneath, not on top." โ The most dangerous emotions in a vocal performance are the ones controlled, not expressed.
- "No breath until this word." โ The absence of breath before a line creates tension that no amount of melody can replicate.
- "Drop voice on this word โ almost spoken." โ Where the melody disappears and only the truth of the words remains.
- "Feel > Perfect." โ The most important instruction on any vocal direction sheet I have ever written.
For the duets "Naa Naa Karke Bhi," "Closer Than We Should Be," "Jeet Ya Haar" โ this direction doubles. Two voices, two emotional arcs, two sets of breath choreography, happening simultaneously and in response to each other. The silence between lines is directed as carefully as the lines themselves.
But the duets introduced a problem that no amount of written direction could fully solve: AI does not reliably recognize which singer should perform which line.
Even when every single line is clearly tagged โ [Male Vocal] for Jimmy, [Female Vocal] for Sofie โ the AI frequently ignores the instruction. Here is exactly what happens: Jimmy sings his line. Sofie is supposed to sing the next line. But Sofie laughs at the end of Jimmy's line โ a reaction, not a vocal turn. The AI counts that laugh as Sofie's turn being over. So the next line, which is written for Sofie and contains clearly feminine words like "karungi," "chahti," "mujhe" โ words that every Hindi speaker immediately recognizes as a woman speaking โ gets sung by Jimmy. A man singing lines written in the feminine first person. The entire emotional logic of the song collapses in one wrong assignment.
A laugh, a gasp, a breath, a moment of silence โ any of these can break the pattern the AI is trying to follow. And it is not just laughs. If there is any overlap between where one voice ends and another begins, the AI resets its own count. It has no understanding of who the character is or what gender the language implies. It only follows pattern โ and the moment the pattern has any ambiguity, it guesses wrong.
And it does not stop at one wrong line. Once the AI misassigns one line, the entire alternating pattern flips. Every line after that point is now wrong โ Sofie's lines are sung by Jimmy, Jimmy's lines are sung by Sofie. The whole song becomes a mirror image of what it should be. A song about a man's certainty is now being sung in a woman's voice. A song about a woman's vulnerability is now being delivered by a man. The characters have swapped bodies, and the emotional truth of every line is destroyed.
Getting the duets right required regenerating each song 20 to 30 times. Not adjusting. Not tweaking. Regenerating completely โ because a wrong voice assignment on one line means everything after it is wrong, and there is no way to fix one line without rebuilding the whole song.
I also needed specific sounds that most songs never require โ a laugh at exactly the right moment, a sharp gasp mid-line, the sound of someone catching themselves before they say something they shouldn't. These sounds exist in the AI's capability. Getting them to land at exactly the right word, in exactly the right emotional register, without breaking the alternating voice pattern โ that took patience and repetition that cannot be automated.
This is not pressing a button. This is directing a performance from the inside โ knowing exactly what the emotion requires and giving the voice the precise instruction it needs to carry it.
Every Instrument Is a Choice. So Is Every Silence.
I choose every instrument in every song manually. Not by accepting defaults. Not by selecting a genre template. By asking โ what does this specific emotional moment actually need?
The instrument selection process has two lists. What goes in. And equally important โ what stays out.
The "what stays out" list is often more revealing than the "what goes in" list. For an intimate song about longing, leaving out a snare drum is a creative decision as deliberate as including a piano. The snare pulls the energy toward pop. The absence of it keeps the listener inside the emotion.
For every song I ask: what does silence do in this moment? What does each instrument add, and what does it take away? I audition combinations until I find the palette that serves the vocal performance rather than competing with it.
The songs that hit hardest in this collection are often the sparsest. The fewest instruments. The most space around the voice. This is not simplicity โ it is precision. Knowing what to remove is harder than knowing what to add.
What AI Does and What It Doesn't
People hear "AI music" and assume the process is: type a description, receive a song. I understand why they assume this. Some people do exactly that. I do not.
In my process, AI is the instrument. I am the musician.
I write every lyric. I choose every instrument. I direct every vocal performance โ the emotion, the breath, the restraint, the crack in the voice, the held note, the spoken word that is almost not a word at all. I shape every arrangement. I make every decision about what stays in and what stays out.
What AI does is execute those decisions at a level of sonic quality that would otherwise require a full production team. It does not replace the artistry. It gives the artistry a production environment. The difference is everything.
The songs exist because of years of writing, hundreds of rejected lines, thousands of small decisions made one at a time over hours and days. AI made none of those decisions. I did.
If you listen to these songs and feel something โ that feeling came from a human being who worked very hard to put it there.
Every cover image you see for this soundtrack was also created with AI โ and every single one required multiple attempts before it was usable.
AI image generation has a well-known problem with human bodies. Hands with six fingers. Arms that bend the wrong way. A shoulder that attaches to nothing. Eyes that are not quite level. Hair that disappears into a wall. Two people standing together where one of them has a leg that belongs to neither of them.
I inspect every image before it is used. Not just for aesthetics โ for basic physical coherence. Does this person have the right number of fingers? Are both their arms attached correctly? Is the hand reaching toward someone actually connected to a wrist? These are questions that should not need to be asked about a photograph of a human being, and yet they must be asked every single time.
Some images were regenerated five times. Some ten. Some more. Each regeneration produces something different โ different lighting, different expression, different body position โ and the inspection starts again. An image that is perfect in every other way gets rejected because one hand has an extra joint or one figure's foot is pointing in an anatomically impossible direction.
What you see in these covers is the version that passed. Behind each one are many versions that did not.