Raian PollockRaian Pollock

Two people,
seventy agents

Flip Education is an AI co-teacher, live in 18 countries and 11 languages, built and run by two people plus a fleet of AI agents doing the work a company normally hires for. Here's how it actually worked, what broke, and what carries over to any other company.

Raian Pollock · fractional COO · about an 18 minute read

01The dinner answer

It comes up at every dinner. Two people ran a product that teachers use in eighteen countries and eleven languages, and everyone wants to know how.

The answer has two halves and the order matters. First half is AI. Second half is the fourteen years I spent building and scaling companies before I ever touched it. Give the same AI to someone starting out today and it looks very different, because AI amplifies whatever judgment is already there. At Flip that meant my co-founder Adriana's twenty-plus years running an education NGO, on the teaching side, and my years of commercial and operations leadership on everything else.

In practice that meant work a team of fifteen to twenty would normally spend a year on, done by two of us in a couple of months. That estimate is mine, and I show the working further down so you can argue with it.

02The bet

I've now done this in both directions. At SEOmonitor, where I was Chief Revenue and Operations Officer, we modernized a running company with live customers, existing systems and a team with habits. That works, and it's careful work, because you're changing the engine while the car is moving. Flip was the rarer case. A blank page, nothing legacy to work around. So we made the bet on day one. The company would be AI-native from the first hour, and we'd only hire a human for something once we'd proven an agent couldn't carry it.

The blank page isn't the point of the story though. It's just where I learned what the destination looks like, faster and with fewer things in the way. Most companies I work with now are the SEOmonitor case rather than the Flip case, and the map still transfers. It's easier to guide a renovation when you've built the house yourself.

One thing surprised me, and it still surprises people I tell it to. Getting everything set up and working on day one was easy. What took months was making it work reliably in production, week after week, with nobody watching every step. Almost everything worth knowing here is about those months.

03Two weeks to launch

The first agent that mattered was a researcher. Mapping national curricula country by country, reading the competitive landscape, helping me put the business plan together. I wanted a clear strategy before anything got built, and that kind of deep structured reading is what AI is genuinely incredible at. The curriculum research went on to become the bread and butter of the product.

Then the building started, and the pace still reads strange to people who lived through conventional product development. The product's history logs every shipped change with a timestamp, so none of what follows is from memory.

Day 1blank page,an evening Day 2curricula for7 countries Day 5Latin America,10 countries Day 818 countries,11 languages Day 1133,000 curriculumpages online Day 13public launch,localized signup
The first fortnight, reconstructed from the product's own change history. Two days after launch, the site carried 142 articles translated into 11 languages.

Over the company's first six months the pace held at roughly 90 shipped changes to the product every week. The number I find most telling, because no pitch deck can fake it, is that 84% of those changes were made with AI in the loop and recorded that way in the history itself. The company wasn't using AI on the side, AI was how the company worked.

Sep Oct Nov Dec Jan Feb Mar Apr May Jun Jul Aug Feb 12: Flip's first day
My actual work record for the twelve months to August 2026, drawn from my public GitHub profile. One square per day, darker means more shipped. 4,226 contributions in the year, the vast majority of them Flip. The wall starts the day the company did.
13 daysfrom blank page to public launch
18 / 11countries and languages at launch
84%of all product changes made with AI in the loop
0 to 17Korganic visits a month, inside four months

04The org chart that isn't

People imagine one clever chatbot. What we actually had looked more like an organization, in three layers.

The first layer was me and Adriana working with AI, live. Product and engineering happened in interactive working sessions, me directing and AI building, all of it running on the written playbooks and standards we developed as we went. The second layer was the fleet, over seventy AI agents with a defined job each, a schedule, and rules about what they could and couldn't touch, quietly running the company's recurring work in the background. The third layer was the scaffolding itself, the written rules and quality bars and institutional memory that kept the first two layers behaving consistently. That third layer turned out to be the asset.

What this normally takes

  • Marketing manager + email specialist
  • SEO lead + content team
  • Localization team, 11 languages
  • Data analyst + BI reporting
  • Finance / bookkeeping
  • Customer support team
  • Compliance counsel
  • Engineers + QA
  • Ops manager holding it together

What ran it at Flip

  • An email guardian that audited every outgoing email for quality, deliverability and brand, every six hours
  • A search-traffic watcher that caught ranking drops daily and wrote a weekly analysis
  • A finance agent that collected spend across ~19 providers every morning, plus a separate one reading invoice emails
  • A business analyst that kept its own investigation journal and produced weekly and quarterly reviews
  • A legal agent maintaining a monthly risk register across four privacy and AI regulations
  • Two support watchers, one fact-checking our support bot's answers, the other hunting for customers who wrote in and got silence
  • And a management layer of agents that reviewed, fact-checked and repaired the other agents

That last line is the one I care most about, because the fleet had middle management. One agent's whole job was fact-checking the reports the other agents sent us, against the raw data, before we read them. Another investigated anomalies on a fixed budget and proposed repairs. Once a month a manager agent reviewed how every agent in the fleet had performed, and we deliberately forbade it from changing anything itself. It could only recommend, and the decisions stayed with us.

The Flip Education fleet map: data sources on the left, collector agents next to them, synthesizer agents, and on the right the lines from all of them converging into inboxes open the full map
The map we actually used. 20 data sources on the left, 51 collector agents gathering from them, 7 synthesizers rolling everything into bigger pictures, and on the far right the only output that ever reached a human, which was 3 inboxes. Open the real thing and filter it yourself.

One detail usually surprises people. Nearly half the fleet used no AI at all. Where a job was pure arithmetic or checking it ran as plain, predictable software, because a calculator shouldn't improvise. Choosing where intelligence is not needed is as much a part of the design as choosing where it is.

Some of these agents are public now. The code is on GitHub if you'd rather read it than take my word for it.

05How anyone found us

None of the above matters if nobody arrives. Flip had no marketing team, no sales team and no ad budget, so demand got built the same way as everything else here, two people and a lot of written rules and a fleet doing the recurring work. Two channels carried it. Search brought strangers in, email decided whether they came back.

The search idea wasn't clever, it was just specific. Teachers don't search for lesson planning software. They search for the exact thing they have to teach on Tuesday morning, one topic in one grade in their country's curriculum in their own language. So we made a page for each one, across eighteen countries and eleven languages, every subject and every grade in each national curriculum. Nobody could staff that as a content calendar. It's a data problem, and it only became possible because the company already had to model those curricula properly to build the product at all. The marketing surface fell out of the product's own foundations.

It worked faster than I expected it to.

Google Search Console performance report for flipeducation.ai showing 37.3 thousand clicks and 2.48 million impressions, with both lines climbing steadily from mid March to late May 2026
The search performance report for flipeducation.ai, 12 March to 22 May 2026. A Flip page appeared in somebody's search results 2.48 million times in that window and was clicked 37,300 times, from a site that had gone live on 12 February. May alone brought 17,310 visits from search. Our own analytics counted 17,742 over the same period, which is about as close as two measurement systems ever agree.

The build produced about 213,000 of those pages. In the window above, 81,762 of them were put in front of a human at least once and 23,873 earned a click. The middle number is the one nobody quotes and the only one worth much, because a page nobody ever sees is a row in a database, not a page.

Building them was easy. Keeping tens of thousands of them healthy normally needs a team, and that's where the fleet earned its keep. One agent told Google about new pages within minutes of publishing rather than waiting around to be found. Another compared each day's rankings against the site's own normal range and only spoke up when a move fell genuinely outside it. That's the difference between a useful alert and a daily email everyone learns to ignore. A third put our own server logs next to Google's records, section by section, so we could answer whether Google was actually visiting these pages at all. Others swept weekly for pages going stale and links quietly going dead.

My favourite one joined what people searched for to what those same people did once they arrived, grouped by the kind of thing they'd been looking for. Not which pages got traffic, but which kinds of searches turned into somebody actually using the product. Traffic is a vanity number until you know which slice of it converts.

Most search reporting tells you what happened. I wanted the boring version. Is anything different today, does it matter, what do we do about it.

One agent turned out to be good at something I wouldn't have guessed an agent could do, which is asking strangers for a favour. Where you rank in search depends partly on how many other credible sites link to you, and the normal way to earn those links is a person emailing editors one at a time for months. We gave that job to the fleet instead. It went hunting for publications writing about the things Flip is about, found the places where their links had gone dead or their statistics were years out of date, wrote a specific pitch for each one instead of a template, and kept track of every thread. It emailed 1,774 people that way, and 8.1% of them wrote back, which anyone who's run cold outreach will recognise as a good rate. The threads it opened ran through education press in the United States, Italy, France and Quebec, each in its own language. Ahrefs, the tool most people use to score how much standing a site has built up, rates domains from 0 to 100 and had us at 15 by August, for a domain that went live on 12 February, with nobody on staff doing press and not one of those emails written by either of us.

Publications researched 1,997
Pitched89% of those researched 1,774
A person wrote back8.1% of those pitched 143
Where the outreach went, from research to reply. The reply rate is counted per person emailed, not per email sent.

Email was the second channel, and the less glamorous one. By the end there were around 50 different emails the company could send, in eleven languages, and almost none of them were newsletters. Each one was tied to something a teacher had done. A welcome sequence. A nudge if somebody signed up and never made their first set of materials, a different one if they'd been active and went quiet, a monthly fresh start, a win back, a note when the thing they asked for was ready. Plus the ordinary transactional mail every product needs to send and usually sends badly.

None of which counts for anything if it lands in spam, so the fleet had what amounts to a deliverability desk. Every morning it sent a real email through the real delivery path to a test address, purely to prove the pipe still worked, because the day that breaks silently is the day you stop hearing from customers and don't know it. Every morning it also checked our sending domains against the industry blocklists, and checked that our mail was still correctly signed so that inbox providers would trust it. Any campaign that crossed three percent bounces paused itself without asking permission. And underneath all of it sat a consent layer every single send had to clear first, so "are we actually allowed to email this person" got answered by the software instead of by somebody's memory.

Both channels tied back to money the same way. Every Monday an agent rolled the whole funnel, visit to signup to first material made to coming back to paying, into one view whose only job was naming the tightest constraint that week. Another watched our experiments and, more usefully, checked whether each experiment was even valid, including whether the traffic split itself had gone skewed.

None of that was a marketing department. It was a set of written rules and a fleet running them every morning, the same answer as everywhere else in this story. The harder question is how you get to the point of letting it run at all.

06Learning to trust it

Nobody sane hands a fleet of software agents the power to spend money and email customers on day one. Most nights I slept fine. There were one or two where I didn't. We built the trust the way you'd build it with a new hire, gradually and with checkpoints.

It started with low-stakes tools and a hard monthly spending cap, so the worst case was the price of a dinner. Then each responsibility got a review rule. The agent does the work, a check reviews it before it goes out, and the training wheels only come off after the reviews come back clean enough times. Underneath all of that sat 3 hard lines no agent could cross alone, whatever the circumstances. Sending external email, moving money, touching anything a customer sees. Those always needed sign-off, by design.

Nobody wrote the standards into a handbook and hoped. We built them into the tools, which refuse to run until the checks pass.

That distinction is the one thing I'd tattoo somewhere visible. Asking AI to follow instructions works about as reliably as asking people to follow instructions, which is to say sometimes. Making the instructions structural, so the system physically won't proceed until the quality evidence exists, is what turned impressive demos into an operation I could sleep on. By the end, the company's shipping pipeline ran more than 50 separate automatic checks before any change could reach customers, and every one of them existed because something specific had once gone wrong.

Which leaves an obvious hole in the arrangement. Plenty of those checks are themselves an AI making a judgment call, and a checker with bad judgment is worse than no checker at all, because you believe it. So before a model got the job we made it sit an exam.

We took real pages off the site and broke them on purpose, in seven specific ways we'd seen things go wrong. Text left in the wrong language, the company name written incorrectly, a missing image, wrong colours, a broken layout, mangled accented characters, content written for the wrong country. We did it across all eleven languages we publish in. Then we shuffled the broken pages in with twenty that were perfectly fine and asked each candidate model what was wrong with each one, without telling it which was which.

The half we expected to be the test turned out to tell us nothing. Every model caught every planted fault, 28 out of 28, all six of them. The entire difference showed up on the pages that were fine.

Gemini 2.5 Flash Lite7$0.13
GPT-4o mini12$0.15
Gemini 3 Flash14$1.15
DeepSeek V3.215$0.28
Llama 4 Scout17$0.21
The same exam, five models that bill per use. Every one of them caught all 28 planted faults; what separated them was how often they cried wolf. Each row is the same 20 pages with nothing wrong with them, and the filled squares are the ones it complained about anyway. A sixth model, running on a flat monthly subscription rather than per use, matched the calmest of these.

That gap is the whole decision. A checker that invents problems gets ignored inside a week, and an ignored checker is a slower way of having no checker. The calmest model in the test also happened to be the cheapest, which was luck rather than judgment. The uncomfortable one was Gemini 3 Flash, already doing this job in production, picked a while earlier because it looked like a sensible choice and never tested against anything. It cried wolf twice as often as the cheapest model on the list, and cost 9x more to do it. Watching it work would never have told us, because a checker that complains too much looks exactly like a checker earning its keep.

Swapping it changed something bigger than the invoice. At $0.13 per thousand pages we could stop checking a sample and start checking every page, in every language, on every change. Sampling exists because checking costs money. Make checking cheap enough and there's nothing left to sample.

07When it broke

Any operations story with only victories in it is a brochure. Things broke here. What running this for months taught me is that the failures were rarely dramatic, they were quiet, and quiet is the expensive kind. The 4 below, told the way they happened.

The watcher that went blind

For 10 days, every new signup was silently left off our mailing list. No error, no alarm, nothing visibly wrong. We found it, fixed it, and did the right next thing, which was to build a watcher whose only job was checking daily that this exact failure never happened again. Then the watcher itself went blind for a week, thanks to an invisible technical fault, and it dutifully reported nothing the whole time. The day we restored its eyesight it caught the exact failure it was built to catch, happening again. Which is when it got uncomfortable. Our original fix had never worked in production. It had looked correct in every test we ran.

So "fixed" stopped being a claim around here and became a measurement. Things that break have to announce themselves, watchers have to be watched, and a repair only counts once you've verified its effect where the customers actually are. Failures must be loud became a company rule that day, and it's now one of the first things I install anywhere.

The bot war

Someone found the free, no-signup version of our lesson generator and ran an industrial operation against it, rotating addresses and automated browsers, at peak costing us roughly 330x what a normal day of AI usage cost. That traffic burned through about 2 years' worth of normal usage in 6 days. Our agents flagged the anomaly quickly and then misdiagnosed it twice before forensic work established what it actually was. A later wave was clever enough to fool our analytics filters into counting it as real visitors, and the automated quality-checker rationalized what it saw instead of doubting it.

Analytics lie unless something is checking them against reality. That's what I'd underline. Every wave ended the same way, with a new alarm that catches that whole class of attack forever and needs nobody to remember to look. The finished version caught a returning attacker within hours, on its own.

The prophecy

One of our quality agents, reviewing an unrelated email bug, wrote down a prediction. Somewhere, a teacher was eventually going to get our emails in English instead of their own language, because of how one line of code stored their country. 24 hours later a teacher in Chile got exactly that. The fix was one line, and one line wasn't enough. 10 days later the identical mistake shipped again in a different email pipeline, because we'd repaired the one place it bit us and left the pattern alive everywhere else.

Fix classes, not instances. The repair that held was a check scanning every email pipeline, present and future, for that whole category of bug, plus a live alarm if a wrong-language email ever goes out again.

The expensive shortcut

Crunching a large public dataset for the research pages we publish on the marketing site, I handed the job to a cheaper AI model instead of the ones we'd standardized on. The written spending rules were right there in its instructions, including a hard one about staying inside the data provider's free allowance. It ignored them inside the first minute and queried the full dataset over and over, all day. The bill came to roughly 30x what even a careless version of the job should have cost. The job itself, for perspective, cost almost nothing to build.

Which model you use turns out to be a governance decision rather than a shopping one. Rules only bind models disciplined enough to follow them, and the discount on the price tag comes back many times over on the bill. After that day, spending power came with a pre-approval gate built into the tools themselves, for every model, however trustworthy.

08A day in the life

The company's day never started, because it never stopped. Here's a shortened tour of what ran while we slept and worked.

  • the night analyst summarizes everything that shipped yesterday
  • a librarian agent files every team report into the company wiki
  • a briefing agent rewrites the shared context every other agent reads, so the whole fleet wakes up knowing what the company knows
  • the audit shift, which collects spending across ~19 providers, checks email reputation, sweeps the business metrics for anomalies and sends a repair agent after anything odd
  • the morning brief lands in our inboxes, the summary a chief of staff would hand a CEO
  • the customer-facing watchers, fact-checking the support bot's answers every half hour and hunting silent failures hourly
  • the fleet fact-checks its own emails to us against the raw data, and files corrections
  • the legal agent reviews the risk register. Yes, the lawyer works Sundays.

Two humans sat on top of all that, reading a handful of emails a day the fleet had prepared, and spending our actual hours on the things that stayed genuinely human. Strategy and product judgment, and the relationships with teachers and schools.

09Keeping the bill boring

Every founder asks what all this costs to run. The number itself is the least useful thing I could tell you, because it swings enormously with company size and ambition, and what Flip spent tells you nothing about what your company would. What transfers is the discipline that kept the bill boring.

Spending becomes visible every morning. Every provider's costs land in one ledger, each number labeled by how trustworthy its source is, so surprises died the same morning they were born. Spending gets approved before it happens instead of discovered on an invoice. Any sizeable AI job has to state its expected cost, justify why it's needed at all, and have an approval recorded before the tools will run it. That gate exists because our early estimates were reliably wrong, in one stretch by 3x, four incidents in a row. And the system checks its own honesty. After every big job the actual spend gets measured against the estimate, and an agent that keeps underestimating loses the right to run big jobs until somebody recalibrates it.

The point is for the bill to stay boring. The bot war and the expensive shortcut above were the two days it got interesting, and both of them bought permanent machinery that kept it quiet afterwards.

Which leaves the estimate I owe you from the first page. The useful number isn't what Flip spent, it's what the same output costs in people. Take the curriculum on its own, roughly 135,000 individual topics, each one read out of a national syllabus in its own language and placed against its own grade and subject. At the rate a curriculum specialist can do that carefully, that one piece of it is ten to fifteen person-years. Add the product, the 213,000 pages, the upkeep in eleven languages, the fifty emails and the measurement around all of it, then cost the lot at what those specialists, engineers and translators actually earn in the countries we were building for, and you land on a team of fifteen to twenty for a year, comfortably over a million dollars. Two of us did it in a couple of months. That gap is the part worth looking at, not the invoice.

10What you keep

Flip proved something I now stake my work on. The value of this operating model isn't the agents at all. Agents are replaceable and the models improve every quarter. The value sits in the encoded judgment, the written rules and quality bars and approval lines and institutional memory that make any capable AI behave like a disciplined member of your specific company. I spent six months building that scaffolding from scratch. It transfers in weeks now, because the thinking is already done and only the fitting is left.

That's what I mean when I tell clients they keep a copy of the operator. The engagement doesn't end with a slide deck about AI strategy. It ends with your company running differently, and the operating system staying behind, still learning, still working after I step back.

The concrete version of that is the catalog, department by department. Every job the fleet did at Flip, and the ones I build for clients now.

For a company that can actually move, month one looks like this. The AI properly set up, access granted, everyone willing to get out of its way, and one real process, or a whole function's recurring work, standing up in production with the quality gates around it. There are two disqualifiers and I name them early. If company policy locks you into weaker AI tools, or the data at the heart of the work can never be shared with AI at all, this model can't deliver its value and I'll say so in the first conversation.

One last thing the story should own. When Flip shifted into a quieter phase, pilot programs instead of full growth mode, the fleet scaled itself down to a skeleton crew and costs followed it down to nearly nothing. No layoffs, no severance, nothing to repair afterwards, just capacity that breathes with the business. Try that with headcount. Flip's own market moves at education's pace and that fight continues on its own timeline. The operating model is the part that already works, and your company can have it.

The audit is where every engagement starts

Two to three weeks inside your operations. You get the savings number, the redesign map and the first encoded playbook, whether or not we carry on afterwards.

Book an intro call