Bridging the Skill Gap between Digital and Physical AI
Bridging the skill gap between digital and physical AI requires solving how we unlock, canonize, and act on our own personal data.
2026 Founder Memo
I’ve spent years launching category-defining products alongside great engineers, C-level positions, product roles, building for consumer data. I led the first consumer robotics, drone and video-recording eyewear initiatives, where I learned firsthand the tension between building tools and building toys, and the challenge of delivering meaningful impact at scale.
But I also collected a unique insight across those experiences:
innovation and adoption usually arrive inside the Trojan horse of fun. The tools that change behavior rarely look like tools at first. They look like play.
The Bridge
I believe we are entering a world where hyper-personalized applications cannot reach their potential until we solve a fundamental problem: how do we unlock, canonize, and act upon our own personal data?
Right now, we manage LLM outputs. In the future, we will manage one AI orchestrator that coordinates a fleet of personal and work-related AIs on our behalf. But that future requires bridging a gap that remains wide open: digital AI and physical AI have no connective tissue. AI assistants will attempt to intercept robotics. Robotics will attempt to learn how humans perform their tasks. But there is no shared infrastructure between these worlds—no common language, no common data layer, no common understanding of who you are and how you live.
Everything I am building toward lives in this gap. An entity focused on hyper-personalized AI that brings physical and digital worlds together.
The “do it like me” company.
Assumptions about where this is heading
Our daily personal and professional lives will require the management of an org chart full of AI agents. There will be conductors at the top of these charts that handle coordination. These conductors will converse with other specialized agents (agent-to-agent) online to curate a specialized version of the entire internet for us—matchmakers of data. In many cases, they will be equipped to make decisions for us and handle menial tasks.
Our primary hardware device will be populated with hyper-personalized apps built specifically for us. Robots may or may not have the onboarding context they need to create $20,000 worth of value from day one. And the Pareto principles in human motivation remain consistent: we all want to improve our aesthetics, longevity, energy, and personal finances while mitigating loneliness.
Personal data: how we got here
When ChatGPT launched, 70% of usage was personal—not work. AI’s first breakout was not productivity. It was a personal cognitive companion. Rewriting messages. Brainstorming ideas. Reasoning through problems. The next wave moved from conversation to execution. Agentic platforms emerged—Manus, Perplexity, OpenAI browsing—showing AI could navigate systems and complete multi-step tasks. Verticalized products like Cursor, Harvey, and Ambience embedded AI directly into domain-specific workflows. These systems work because they maintain state, understand how work is done, and operate inside existing processes.
But here is the gap: although AI’s breakout was personal, nearly all subsequent infrastructure has been built for work.
This is not an accident. It is structural. Workflows are explicit. Outcomes are measurable. Context is structured and permissioned. Modern work tools are cloud-native, exposing canonical APIs and schemas. AI tooling evolved accordingly—optimized for desktop environments with native integrations and protocols like MCPs.
Personal life breaks every one of these assumptions. Most people live on mobile, where context is siloed across apps with limited or nonexistent APIs. Personal life does not run on workflows or maintain a single persona. Models are trained to sound right, not to be accountable. And until that changes, AI will remain infinitely helpful in theory and unreliable in practice.
What I have built to test these assumptions
I have been personally vibe coding to live inside these gaps.
A stateful health agent I call web.ME—built on my 23andMe data, Apple Health, meal photos, and peptide schedules using Letta memGPT—that lets me query my own body instead of WebMD.
A meal prep workflow spanning Memories.ai, Apple Notes, ChatGPT, OCR screenshots of my pantry, and two separate Instacart lists optimized for my high-protein diet and my sixteen-year-old’s gluten-free needs.
A closet reconstruction project that scrapes email order confirmations, remakes images via Google Whisk for resale, and connects to a listing agent built with Claude Code, Google AI Studio and python on desktop.
A WhatsApp travel agent prototype that aggregates group chat inspiration into shared Google Maps lists and babysits reservation openings built using Manus.im.
A fitness PR (personal record) dashboard correlating Garmin, Oura, and Apple Sleep data with weather, course terrain, and time of day to help my son find the conditions where he runs his fastest cross country & track times built with an MCP server through Vercel and Figma AI Make for UI.
A co-parenting dashboard that filters school emails, tracks grades and homework, and scrapes the high school calendar so both households stay aligned via Claude projects and Cowork.
Every project hit the same walls: siloed data, no persistence between sessions, tools that assume a single user on a desktop, and zero infrastructure for multi-source, multi-persona, action-oriented workflows. I did not set out to validate a thesis. I set out to solve my own problems. The thesis emerged from the friction.
The “Do It Like Me” Problem
Everyone talks about personalization. I am talking about something different: the gap between how a system performs a task generically and how *you perform it.*
Robots are coming to our homes, but not to take over work tasks. They are coming to help with physical, personal tasks: laundry, cooking, organizing. I am not convinced these robots will arrive equipped with skills for the actual demographics they serve. Will they know how to fold women’s intimates? Hang delicate items? Fold small baby clothes? And even if they can fold a standard towel, can they learn to do it like me?
This is not a preference setting. This is an entire training problem that nobody has infrastructure for. The resolution of personalization required for physical-world tasks is orders of magnitude higher than what anyone is building for, and we have no capture mechanism for it.
So how do we train AI—and eventually robots—on the personal, idiosyncratic ways we do things? Not the generic way. Not the average way. Our way.
There is a concept I keep returning to: a personal Montessori for robots. But school does not begin when the robot arrives. It begins years before—in a virtual world where you are already teaching the AI who you are, without realizing that is what you are doing.
We will not train robots through programming. We will train them through living—through a pipeline that starts in play and ends in physical competence.
Phase 1 — Years of virtual data, collected through play
Could the AI learn all of this through play—through a game you actually enjoy? What if we had a way to simulate real life—Minecraft for adults, The Sims reimagined?
Consider the grocery problem. Most humans do not trust other humans to choose produce the way they would. They consider themselves picky about ripeness, firmness, blemishes. Others cannot be bothered by the overwhelming task of adding twenty-five items to a cart while mentally inventorying what is already in the pantry. If you’re open to using the Instacart app via chatGPT, there is no memory for dietary restrictions, your family’s allergies, your weight loss goals. But before you even get started, you’ll have to re-watch, transcribe and screenshot your inspiration for meal prep from TikTok or Instagram.
What if you could teach your AI what you look for in an avocado? or not buy strawberries in a plastic clamshell?
What if you could wander leisurely through a virtual world that reflected your interests—into book-tok, into meal prep, into the closet of your dreams in your New York apartment, into a beautifully organized pantry that looks like yours could look? Not an endless scroll. A space. A place you could walk through, explore, and learn from.
In the virtual world, your bathroom is populated with all of your skincare, beauty, and haircare products—scraped from your email inbox receipts. You add inspiration from TikTok tutorials. You get step-by-step instructions with keyframes and an AI voiceover. Your kitchen reflects your actual pantry. Your closet holds your actual wardrobe. Every interaction is data—preference data, taste data, behavioral data—captured over years, passively, through a medium you actually enjoy using.
This is not a simulation for its own sake. This is the training set. Years of virtual living that teaches the AI your aesthetics, your routines, your tolerances, your idiosyncrasies—long before any robot is involved.
Phase 2 — The physical learning environment
Then the cameras arrive. A six-month onboarding period where sensors throughout the home capture how you actually perform tasks in three dimensions—feeding this into a neural network that closes the gap between what the virtual data predicted and how you actually move, fold, cook, organize in physical space.
The virtual world taught the AI what you want. The Montessori period teaches it how you do it. The delta between these two is where the real personalization lives—the difference between knowing you prefer your towels folded in thirds and knowing the exact pressure, speed, and tuck you use when you fold them.
Phase 3 — The robot arrives equipped
The robot does not arrive cold. It arrives with years of context from the virtual world and months of physical calibration from the Montessori period. It knows both what you want and how you do it.
Perhaps it comes equipped with a “beauty package” that can be downloaded and personalized—trained on your saved tutorials, your product inventory, your morning routine. It references how you actually apply your makeup. It practices dexterity with brushes and small compacts while you are at work. And then you sit while it gives you a makeover each morning.
It observes how you cook and what each family member eats. It calculates calories consumed. It knows what is left in the pantry, the spice rack, the refrigerator. It aggregates data from your online grocery order, your Kroger curbside pickup, and your analog trip to Trader Joe’s.
This is the pipeline: virtual data over years, physical calibration over months, and a robot that arrives not as a blank slate but as something that already understands your life. The training began the moment you started playing.
The Data You Need to Make It Work
The moment you try to build a personal app around your actual life—meal planning, home design, virtual closet, fitness, beauty—you run headlong into a structural problem no one has solved. You need three kinds of data. They do not play well together.
Owned data is yours.
Your health metrics, photos and screenshots, personal messages, notes - it lives on your phone. There are no IP questions here. This is yours to use—if you can collect it and understand it.
Curated data is where it gets complicated.
The last few weeks of saved TikToks and Instagram stories. The recipe videos you bookmarked. The design inspiration you screenshotted. You did not create this content. You selected it. The curation is yours—the taste graph, the collection logic, the act of saying this, not that. But the underlying content belongs to someone else.
Derived data is where it gets dangerous.
This is what the LLM generates when it remixes all of the above into a meal plan, a shopping list, a design mood board. It is a blend of your taste graph, someone else’s intellectual property, and a model’s interpolation. Who owns that output? Who gets paid to create it, distribute it?
*Robots, AIs and agents all have the same problem. They all need context protocol—a standardized way of saying: here is what I know, here is where I learned it, and here is what is mine to share and what is not.*
The Taste Graph as Asset
First, we will create for ourselves—to organize our world. Second, we will share with friends—to send our digital AIs to do their bidding, equipped with our knowledge. And then, we will seek to make money from the time we have spent collecting, curating, and creating. This is the natural sequence, and each stage introduces a harder problem than the last.
My friend Courtney has collected the best recipes. My friend Theresa has interior design preferences that make you feel like you are living at a resort. My friend Sashia spends her time curating all of the best “girls night out” spots in Los Angeles. Each of them has built—whether they know it or not—a personal knowledge base of extraordinary value.
Stage one is already happening. Vibe coding has made it possible for anyone to prototype a personal app. A meal planner that actually knows your dietary restrictions. A closet organizer built around your wardrobe. A fitness tracker that connects what you eat to how you feel to what you weigh. These are apps built for an audience of one—single-tenancy tools designed around your life, your data, your preferences.
The personal data that powers these apps is different from behavioral data. It is not how you browse or what you click. It is taste data. Preference logic. The invisible rules that govern why your home looks the way it does, why your cheese board delights, why your friends ask you to plan the trip. It lives in the gap between what you save and what you make from it.
*The curation is the creation. The act of collecting and organizing—not producing—creates something worth owning, sharing, and eventually monetizing.*
Can an app travel without its content?
Is it possible to separate what makes an app valuable from the data that makes it personal—and from the content that was never yours to begin with?
Can we share the workflow logic—the structure, the rules, the preference weights without sharing the data itself?
And what happens to the curated layer? If my app is built on top of saved TikTok recipes and bookmarked tutorials, can those travel as references—pointers back to the original creator content, tagged with metadata about why I saved them—rather than as copies?
Then there is the hardest question: what about derived data? If an LLM remixed my owned data, someone else’s recipe, and its own interpolation into a meal plan—what is that output? Can it be stripped and regenerated on import, rebuilt?
Why Seeing Is the Key
We are highly visual creatures. We record, remember, and understand through images. There is an opportunity to bridge the gap between digital and physical AI through vision—starting with the highly visual inputs we already create: shopping screenshots, meal photos, closet pictures, text message threads full of links and images.
Consider: you screenshot an egg.
When we talk about computers “seeing” images, we usually mean identification: can it recognize the objects, can it read the text. But identification is not understanding. A model can recognize an egg. It cannot, from recognition alone, know what to do with that egg.
City Guide: If that egg appears inside a Bon Appétit review titled “What to Order at LA’s Hottest New Openings”—tempura-soft-boiled, nestled in a bowl of tonkotsu at the new place in Silver Lake—the context is discovery. You are bookmarking a restaurant. You want the location, the reservation link, the name of the dish.
Recipe: If that same egg appears inside a YouTube cooking tutorial for homemade ramen, the context is execution. You need the full recipe. You need keyframes showing how long to boil it, when to ice-bath it, how to peel it. You need to know what else to order.
*Same egg. Completely different needs. The egg is not the information—the egg is a pixel cluster. The information is everything around it: the source, the format, the intent, the surrounding text, the genre of the content.*
A city guide and a recipe are different knowledge structures. They require different parsing, different outputs, different downstream actions. This is what gets lost when OCR is an afterthought. We extract the text, we tag the objects, and we throw away the scaffolding. But the scaffolding is the meaning.
Machines learn in zeros and ones. Most of our personal data lives in images. I believe there is something special in unlocking our highly visual personal lives. Something worth building.
What I Am Building Toward
I keep coming back to a strategic focus on action-oriented systems. AI doers. A do engine. Not systems that help you research or plan, but systems that take responsibility for outcomes.
Everything points toward a future where all apps on your phone become apps tailored specifically to you. Where all online systems become standardized systems filtered through your personal agent. Where hardware and software become symbiotic, and you can qualify and disqualify agents or robots based on their skills.
Core questions I am wrestling with
On training: How will we train agents on our personal, real-world tasks? Mimicry worked beautifully for Manus.im’s start as a browser extension—Monica.ai, now exited to Meta. What is the parallel for our locked phones, given that Apple seems unlikely to build this themselves?
On determinism: How deterministic should AI be, and how creative, and when in the timeline of the task? Action-oriented systems require precision. Creative systems require latitude. The answer is probably both, but the boundaries are unclear.
On the art of AI unmaking: Most AIs are trying to make things by understanding what good looks like. But perhaps we need to get really good at unmaking things so we can make them again. Teaching is about breaking something you know how to do into smaller, manageable tasks that you can demonstrate to another human or machine.
On real value: Is AI evolving so fast that it is creating a false momentum of value? If I can quickly build a prototype, does this help me find users any faster? How much money is being dumped into tokens based on curiosity and FOMO but lacks any real output?
On form factor: everything starts with a prompt box today, but is that the right user interface for this kind of experience? We need to test a variety of design paradigms that could include workflows, org charts, virtual worlds, ambient voice-enabled inputs in order to understand how to best achieve the user goals.
How will vibe coding evolve if the tools we have are not built for how we think
Imagine you could prompt an AI to make you a car. And it does. The car is beautiful. You can sit inside it. It even drives—in a straight line. But then you need to make a turn. Or the engine makes a sound. Or you want to add a bike rack. And suddenly you are standing in front of the hood with no idea how to open it.
This is vibe coding today: you can make the car, you cannot fix the car. Engineers know what is under the hood because they put it there. Consumers see the car as a single object. Non-technical users do not think in file structures and dependencies. They think in tasks, outcomes, workflows. Building for them requires a fundamentally different metaphor—perhaps code stops being files and starts being something closer to a conversation, or a recipe, or a set of rules you teach by example.
I've spent a year obsessing over this gap
My Findings in User Research

- 51 calls with friends, family & founders
- 135 survey recipients
- 35 personal apps vibe-coded (see what’s currently possible)
- 80% regularly delay real-world tasks because they feel annoying, confusing, or stressful. Strong indication: This is not a niche or personality issue. It’s a broadly experienced pain.
- 66% prefer full task handoff with confirmation over help figuring out what to do
- 93% say proof-of-completion is absolutely required
- 65% willing to pay $5-25 for reliable task execution This is not seen as a “free AI feature” problem.
Deeper Research
50M+ GPT conversations 23,000+ apps & tools
Abstract Concept: Healing Action Models
HAMs: A research-based solution I’m currently exploring with founder friend Ross Ingram
Conclusion
Almost everything we do begins with media, messaging and notes. The question is not whether we need a bridge between digital and physical AI. The question is who builds it—and whether it arrives before the robots do.