What It Cost to Produce My Own Story
Making And Now? — a six-minute AI-generated film about my own decade — taught me what these tools can hold of a human story, and what they can't.
March 1, 2025·project / ai / film
01 — Why I made it
I have spent years thinking about what information does to a person — how the right image at the right moment can crack something open, show you a version of yourself you had no other way of imagining. So when generative AI tools started making it possible for anyone to produce a film, I didn’t want to observe from the outside. I wanted to get my hands dirty.
But I also had a story I needed to tell.
The film is called And Now? It was born from a moment of pause. As I started my MBA at MIT Sloan, I wanted to look back at how far I had come — from a small town in China to New York City, and now to Cambridge. These years were not a straight path. They were a decade of resilience, discovery, and transformation that existed only in my memory, and that no production company would ever take on. When I saw what these tools could do, I thought: this is the form.
So I had two motivations, and I didn’t try to separate them. I wanted to make something personal, and I wanted to develop real fluency with these tools — not the fluency of reading about them, but the kind that only comes from using them until they frustrate you. Both things happened. They happened at the same time, and they taught me things I couldn’t have learned any other way.
02 — What it cost
The film runs about six minutes and follows me through four chapters, each with its own color. Winter blue for New York beginnings — I arrived in winter, and life felt cold and empty. I had nothing except determination and hope. Spring green for growth and awakening — after three years in a nail salon and community college, I transferred to Baruch, and my world started opening up. Summer gold for energy and momentum — long hours at KPMG, new industries, new people, learning faster than I ever had. Autumn gold for harvest and new beginnings — nine years in, I finally brought stability to my family, and that stability meant it was time for a new chapter.
I knew these colors before I knew anything else about the film. They were how I had always understood my own story — not as a timeline but as a set of feelings, each with a season, each with a temperature. The challenge was translating them into something a machine could use.
The hardest part wasn’t the technology. It was me.
I needed the character to be consistent — the same person moving through different seasons of her life, appearance changing in ways that meant something. The red coat. The shift in outfits and hairstyles and glasses that mirrored my own evolution. I knew exactly what I wanted. I had lived it. And I could not find the words to give it to a machine.
So I prompted. And prompted again. Each attempt came back close but wrong — the wrong coat, the wrong posture, a face that captured maybe 60 to 70 percent of who I was but not the rest. I’d adjust a word, reframe the description, try a different angle. The machine was patient in a way that started to feel like indifference. It would give me anything I asked for. It just couldn’t give me what I meant. Glasses were especially hard — small variations would break continuity across shots, and fixing one thing would shift something else. Each second of the film felt earned through iteration in a way I hadn’t anticipated.
That gap — between what I meant and what I could say — is what exhausted me. I had always assumed my own story was something I possessed fully, something stored and retrievable. Prompting taught me otherwise. I know my life the way I know a face: immediately, completely, without being able to explain how. The moment I had to describe it in language precise enough for a machine to act on, I discovered how much of it lived below words.
This is what I mean by cognitive fatigue. Not tiredness from long hours, though there were long hours. It was the exhaustion of being the translator of your own life — of having to make explicit what you have only ever known implicitly, for something that cannot meet you halfway.
03 — What I learned about the tools
Making the film taught me things I couldn’t have learned by reading about them.
The first is about integration. I built the film across four tools — Midjourney for static master frames, Runway for motion, ElevenLabs for voice and sound, CapCut for final assembly. Moving assets between them, managing files, deciding which tool to reach for at each step — this consumed creative energy that should have gone toward the story. When I later used ElevenLabs Studio, which consolidates several of these steps into one environment, I felt the difference immediately. Creative momentum is fragile. Friction kills it. An integrated platform doesn’t just save time — it keeps you inside the work.
The second is about controllability. Six months ago, re-prompting was often the only way to adjust an output. Now tools like Runway offer backdrop replacement, character swap, color grading, lighting adjustment. This is real progress — it shifts the creator’s role from requesting outputs to directing them. But as features accumulate, platforms become harder to navigate. Knowing which tool to reach for, when, and how different controls interact requires experience and tacit creative knowledge that most people don’t have. The technical barrier to production is lower than ever. The cognitive barrier to effective use is rising. These two things are happening simultaneously, and the gap between them is widening. In practice, feature expansion disproportionately benefits creators who already understand creative workflows. For everyone else — for the person who has a story but not a production background — more features can mean more paralysis. During my own process, momentum depended less on what was possible and more on how quickly I could decide how to work within the system.
The third thing I learned is about animation. When I framed scenes as animated or stylized rather than realistic, everything worked better. Imperfections became aesthetic choices rather than failures. AI-generated visuals have a particular texture of unreality that fights against photorealism but sits naturally inside animation. This isn’t a limitation to engineer around. It’s a signal about where these tools are most honest.
04 — Where this is going
Making the film clarified something I had been thinking about abstractly: production is becoming the solved problem. The harder questions are what comes after.
In the near term, the most effective applications of generative image and video aren’t in open-ended creative tools — they’re in domains where what happens after generation is already defined. Shopping and marketing are the clearest examples. AI shopping agents, avatar-based try-on experiences — these work not because the generation is impressive, but because the generated content is embedded in a decision flow. Users know exactly what to do next. The content is useful before it needs to be meaningful. This contrasts sharply with open-ended tools like Sora, where generation is powerful but the experience often ends at the output. Users create something, briefly admire it, and move on — not because it isn’t good, but because there’s no structure around what it’s for. Production without placement is a dead end.
In the medium and long term, AI’s most significant contribution to entertainment won’t come from replacing existing formats but from enabling new ones. Animation is the leading edge. Its tolerance for stylistic inconsistency, for non-photorealistic motion, for generative artifacts — all of this makes it structurally compatible with what these tools do naturally. AI also compresses the timeline and cost of experimentation dramatically. Storyboarding, visual exploration, mood testing — tasks that once required significant coordination can now be done by individuals or small teams. This doesn’t replace studios. It lowers the threshold for who gets to test an idea.
The long-term opportunity isn’t substitution. It’s expansion. As AI lowers the cost of making ideas visible, more stories become tellable — stories that would never have justified a traditional production budget, told by people who never had access to the tools. I am one of those people. And Now? is one of those stories. The entertainment landscape will be reshaped from the edges rather than the center, and I think that matters.
05 — What the film gave back
When I finally watched the finished film, I recognized myself.
Not in every frame. The models captured maybe 60 to 70 percent of my real appearance — and the rest, the expressions and details that live in the other 30 percent, I eventually stopped trying to recover. Some things didn’t survive the translation. I made peace with that.
But the sound was right. The vibe — the quiet, the sense of someone who has traveled a long distance and is still traveling — that landed. The four colors held. Winter into spring into summer into fall, each chapter turning into the next. Something came through, not despite the translation but through it.
I made this film to understand what it would take for a machine to hold a human story. What I learned is that it can hold parts of one — the parts you can name, the parts that survive conversion to language. The rest stays with you. Maybe that’s always been true of storytelling. Maybe every medium is just a different kind of lossy compression, and the question is only what you’re willing to lose.
The cognitive fatigue I felt is a design problem. The gap between memory and prompt is a product problem. What happens after production — how content finds meaning, finds an audience, finds a place in someone’s life — is a business problem. I don’t think these are separate problems. I think they’re the same problem approached from different angles, and the most interesting work in this space lives exactly at their intersection.
I wear a red coat at the beginning of the film. By the end, I’m not wearing it anymore. I knew what that meant before I could say it. It took me hours to get the machine to show it. When it finally did, I watched it three times.