iOS · TestFlight beta

Split the bills by texting a note.

Nestlit is an expense app for housemates. Type one line, the way you would in a group chat, and it becomes a shared expense. No forms, no keypad.

Got it — here's what I'll add

Now try it yourself on the phone. Type a note, or tap one:

The phone is a web replica with a simplified version of the app's instant parser. Send a note, mark a bill paid, then open the chart tab to see the balances move.

9:41
N
Our Nest
You, Sam & Jordan
Got it — here's what I'll add

The form is the problem

Expense apps ask for an amount, a title, a category, a payer and a split. That is five decisions for one coffee, so people stop logging the small things and the balance stops being true. Housemates already describe expenses to each other in one line. Nestlit makes that line the input.

Typing “dinner 86 sam paid split” shows a preview card: Dinner, split 3 ways, Sam paid, 86
Type a noteSee how it was understood as you type.
The expense appears at the top of the shared feed
It lands in the shared feedEveryone in the house sees it in real time.
Insights screen with who owes whom and this month's spending per housemate
Balances stay trueWho owes whom, and where the month went.

How it works

Two parsers work together: a deterministic one for speed, and an on-device language model for messy phrasing.

1 · AS YOU TYPE

An instant parser reads the note

It finds the amount, payer, split, category and due date on every keystroke and shows a preview card. No network, no waiting.

2 · ON SEND

The expense posts immediately

What you saw in the preview is what gets saved. Adding an expense never waits on the model.

3 · IN THE BACKGROUND

The on-device model refines it

Apple's Foundation Models re-reads the note and patches the row only if it changes something meaningful. If the model is unavailable, nothing is lost.

What the model is trusted with

A language model is good at loose phrasing and unreliable at dates and identity. So each field has an owner, and the model's output is checked before it touches shared money.

FieldOwnerWhy
Due date, recurrenceDeterministicA date detector is more reliable than the model on “due Friday” or “every 2 weeks until Dec”.
PayerModel + codeThe model extracts a name; code matches it against the real household. An unknown name falls back to you, never to a guessed person.
Name, category, splitModelThis is where phrasing actually varies. Category is a fixed list, so the model cannot invent one.
AmountParser firstThe parser handles 2k, 124,50 and “2 pizzas 30”. A non-positive amount from the model is discarded.

Product decisions

The interesting choices were mostly about what the AI should not do, and what the product should not have.

The model never blocks the user

An early build waited for the model before saving, and adds stalled for seconds on some phones. Speed on the core action mattered more than a slightly better first guess.

On-device, not cloud

These are a household's finances. The model runs locally: nothing leaves the phone, there is no per-request cost, and it works offline. The price is device coverage.

Working without AI is the default path

On older phones, or on any model error, the app keeps the parser's result. The product is complete without the model; the model only makes it more forgiving.

Show the interpretation first

A wrong parse is visible in the preview before it becomes a shared record, and every expense stays editable afterwards.

Cut: receipt scanning

Built as a stub, then removed. OCR that guesses wrong erodes trust in a money app, and a camera flow fights the text-first idea.

Cut: payments and trips

Settling up records that a debt was paid; the app never moves money. One household per person, focused on the everyday case.

What real use taught me

Using it in a real household through TestFlight surfaced things no design review did.

  • “Playstation game 15th due 1000” was booked as 15. The parser read the ordinal as the price. Every misparse found in use becomes a permanent test; the parser now has 41.
  • Silence looks like a crash. A note with no amount used to do nothing at all. It now says what is missing.
  • Two phones disagree. A housemate who joined late was missing from balances, so the app said “all settled” while money was owed. Balance correctness across devices went ahead of every new feature.
  • A green test suite is not a working product. One build shipped where every add failed, because no test touched the backend. Tests now run two simulated devices against one fake backend.

What's next

In order of priority.

  1. An eval set for the model path. The parser is well covered; the model's refinements are not measured yet. A labelled set of real notes, scored per field, so a prompt change is judged by numbers.
  2. Measure whether the model earns its place: how often its refinement changes the result, and how often the user then edits it back.
  3. A fallback for phones without Apple Intelligence, weighed against the privacy promise that makes on-device attractive.
  4. Accessibility: Dynamic Type and VoiceOver labels.