Back to Blog
Work & AI•2026-10-08•4 min read

How we trained the best on-device AI for time tracking categorization

Cronus is an AI time tracker. It reads the active app, window title, URL and page text, and files each stretch of your day under one of your categories, like "Coding", "Meetings" or "Distraction". The page text is what tells a pull request review apart from a YouTube rabbit hole in the same browser.

Until this week, that text went to our servers, where Jev, a decision model from TypeSafe, picked the category. Since Cronus 2.2.77, the pick happens on your Mac and the text stays there. That took a model small enough to ship with the app and accurate enough to trust.

What the model has to do

The model gets a description of your activity and your category list, and picks one. We need the pick to be right, fast and come with a confidence score. Cronus shows that score next to each activity ("78% Coding · next: Research") and uses it to decide when a past categorization of the same page can be reused.

How we scored the models

Every model chose categories for 571 activity samples held out from training, using real category lists.

  • Agreement: how often the pick matched the category Cronus already had.
  • Flips: picks on the wrong side of productive and unproductive. Coding filed as "Research" is minor; coding filed as "Distraction" wrecks your score for the day.
  • Speed: median (p50) and p90 time per decision on an M5 Pro.

The reference was Jev 1.13, our cloud model.

Results

ModelDownloadAgreementFlipsp50 / p90
Jev 1.13 (our cloud model)–86.6%2.7%272 / 539 ms
Qwen3.5 0.8B, fine-tuned by us0.54 GB88.7%5.3%218 / 298 ms
Qwen3.5 2B, fine-tuned by us1.31 GB89.7%4.8%401 / 539 ms
Lev (4B)3.0 GB79.6%11.1%2063 / 2915 ms
Kev 4B3.0 GB68.7%17.2%1129 / 1650 ms
Qwen3.5 4B2.7 GB65.5%20.2%1275 / 1980 ms
Apple on-device modelbuilt in64.3%20.2%2087 / 3411 ms
Clef-flash (9B)6.5 GB63.7%22.3%2257 / 3235 ms
Gemma 4 E2B3.3 GB50.4%31.7%512 / 701 ms
MiniCPM5 2B1.56 GB49.2%34.9%509 / 736 ms
Qwen3.5 2B1.28 GB48.1%35.5%424 / 597 ms
Laya0.45 GB46.6%37.2%222 / 329 ms
Prism 0.8B0.53 GB39.9%41.4%1016 / 9075 ms
Julia-10.17 GB35.7%43.1%57 / 84 ms

No general-purpose model passed 66% agreement, even at 4B or 9B parameters. Our fine-tuned 0.8B model matched the cloud model and was the fastest of the accurate ones. It agrees slightly more often than Jev but flips twice as often (5.3% against 2.7%), so we consider the two roughly even.

Why we didn't use Apple's model

macOS 26 ships an on-device model that apps can call through the Foundation Models framework. It's free and needs no download, so we tried it first.

It reached 64.3% agreement with 20.2% flips, and leaned productive: it picked a productive category for 76% of samples where the reference had 65%, so your days would look better than they were. It took 2.1 seconds per decision against our 0.2. It returns only an answer, with no probabilities, so Cronus would lose its confidence scores. And it requires macOS 26 with Apple Intelligence turned on, while our model runs on any Apple Silicon Mac, including 8 GB ones.

Its safety filters never refused a screen, and its 4,096-token context never truncated one.

Fine-tuning a small model for one task

A general model can't know that "Arca Coding" and "Coding" mean different things to you, or that a terminal open on a side project isn't client work. So we fine-tuned Qwen3.5 0.8B on a large set of examples labeled by our cloud model. Each label included the cloud model's confidence in every category, so the small model learned how sure to be as well as what to pick. This is called distillation.

The model doesn't generate text. Each category gets a letter, and one pass through the model gives the probability of every letter, which is why a decision takes about a fifth of a second. We shuffled the category order during training so the model couldn't learn that the first letter is usually right.

One pass over the data, about two hours on a single GPU, raised agreement with the cloud model from 47% to 84%, and to 86% on a second held-out test set. Compressing the weights to 4-bit brought the download to about 540 MB and agreement from 90.8% to 88.7%. A 2B version scored a point higher but was twice as slow and two and a half times the download, so the 0.8B model ships.

What surprised us

Training for the task mattered more than size. Clef-flash, a 9B decision model, scored 63.7%. Lev and Kev, 4B models trained like Jev, beat the plain Qwen3.5 4B, and our 0.8B model beat them all.

The compressed model first scored 61%, as if 4-bit had ruined it. The bug was in our scorer, which counted both "A" and " A" (with a leading space) as A, while training had only shaped the bare letters. Scoring the bare letters gave 88.7%.

Adding hand-corrected categories to the training data made the model slightly worse. Corrections are rare and personal, and a model this small seems to treat them as noise.

What it costs your Mac

The model downloads once, about 540 MB, in the background during onboarding. It uses about 775 MB of memory while loaded and unloads after 10 idle minutes. In the app, a decision takes about half a second, and the model averaged about 1% CPU during normal work.

Trying it

New installs on Apple Silicon categorize on your Mac from the start. Existing users are offered the switch when they update to 2.2.77, or can turn it on in Settings. Cloud categorization is still available there and is about 3x faster end to end; in that mode Cronus removes passwords, keys, card numbers and phone numbers on your Mac before anything is sent.

In on-device mode, our server stores only metadata (app names, window titles, URLs, timestamps and categories) so your dashboards, history and iPhone app stay in sync.

cronus

Find out where today actually went.

Automatic time tracking for Mac and iPhone. On-device AI sorts your day, and your screen never leaves your Mac.

iPhone app