Key Takeaways
- Bloomberg reported that Meta is quietly deploying human contractors to handle live phone calls and concierge tasks for its upcoming Muse personal AI assistant.
- Jordi Hays argues human-backed concierge services cannot scale to Meta's billions of active users without collapsing under labor costs.
- John Coogan explains that human fallback functions as a data collection flywheel, generating proprietary training data on edge cases where current voice and reasoning models fail.
- Startups facing high token costs or brittle agent workflows can treat human fallback as an R&D data investment rather than a permanent operational expense.
The Disagreement
When Bloomberg reported that Meta was using human contractors to take over phone calls and live concierge requests for its new assistant, Muse, the immediate reaction was skepticism. Paying humans to act like software is an old tech punchline. It brings to mind early concierge startups like Magic or Operator that bled cash trying to scale manual labor disguised as automation.
Hays pointed out the arithmetic problem right away. Meta builds software for billions of consumers. If a fraction of those users begin routing phone calls through Muse, the labor expense explodes.
“It feels completely unsustainable for meta to roll out a product to billions of people where users might just be like like if I had like something I could text and say, 'Hey, call this person, call that person,'” Hays argued. He added that the revelation confirms an old suspicion about chat tools: “Yesterday we watched a video where he said, 'I guarantee that there are humans behind the chat apps that you use manipulating the information' and uh he is correct at least in the short term.”
Coogan saw a different play entirely. Meta is not trying to build an outsourcing call center. They are buying the one thing missing from public internet datasets: high-fidelity traces of messy, real-world task execution.
“So Meta is not getting that data, but if they have a human in the loop, the human who does the task that's just beyond what the model's capable of, they're also generating training data because they can probably train on that,” Coogan explained. “So I would view this human in the loop thing more as them doing data collection and and you know creating more training data for them than a permanent solution to the product problems.”
Who's Right (and When They're Wrong)
Hays is right if Meta treats human fallbacks as a product feature. If you sell a user an AI that secretly requires a human operator, your gross margins erode as adoption grows. We are already seeing this pressure across the industry, such as reports of negative gross margins in agentic legal software when token consumption and complex tool calls spiral out of control.
Coogan is right because Meta is treating contractors as an R&D expense. Current voice models struggle with noisy acoustic environments, rude receptionists, IVR phone trees, and ambiguous instructions. Scraping the public web cannot fix these failures because that data does not exist online. When a contractor steps in, answers the restaurant host, deals with background noise, and completes the reservation, Meta captures the exact audio, transcript, and decision path needed to fine-tune future models.
Human-in-the-loop only works as a short-term data engine when you have a strict pipeline to feed edge-case transcripts directly back into model evaluation. If you use human fallbacks without logging, structuring, and training on the failure points, you are just running an expensive concierge desk.
What to Do With This
Audit your product's agent drop-off points this week. Set up an automated routing rule that alerts an engineer or operations specialist whenever your AI assistant fails an API call or hits a confidence score below 70 percent. Have the human resolve the request manually, log the exact input-output pair in a dedicated evaluation dataset, and use those specific failure cases as your fine-tuning benchmark for your next model release.