What needed solving
Agent tools that read documents, write reports and run on a schedule are now common, but most are paid cloud services, and every file they touch leaves the business. Firms that handle drawings, contracts and site reports often can't send those to a third party. Local alternatives exist, but they are hobbyist tools with few guardrails, and small local models are unreliable at using tools: in our first smoke test, 7–8B models failed basic multi-step file tasks.
How we're building it
Simam installs and manages a local model engine itself, then adds 'teammates': named agents with a working folder, one conversation, editable notes memory and an activity log. Each tool (search, read, list, write, run command) is set to Off, Ask or Allow; writes and commands default to Ask, showing a preview and waiting for approval, by notification if the window is closed. Routines run a teammate daily, every N hours or when files in a folder change, from the system tray. Before building features we added a reliability layer: recovery of tool calls written as text, a repeat-call guard, argument checks, one call per round, a 120-second command limit, and a check that catches an answer claiming a file was saved when no write actually happened.
What we learned
We measured the layer with a fixed five-task eval: summarise a site note into a file, find which of three files names the crane operator, read a file from a loosely worded name, count .log files with a shell command, and answer 12×7 without tools. Over 5 runs per task (25 attempts), qwen3:8b scored 23/25 and qwen2.5-coder:7b 20/25. The layer mattered most for qwen2.5-coder, which writes tool calls as text: it scored 3/15 with the layer off and 12/15 with it on, over 3 runs per task. llama3.1:8b stayed below the 80% bar; the app does not suggest it for teammates and warns if it is chosen. Only qwen3:8b completed the multi-file 'find' task. A re-run on Oct 6 after adding the claim check scored qwen3:8b 24/25 and qwen2.5-coder:7b 19, 20 and 20 of 25 over three runs, so no regression. In live use qwen2.5-coder:7b sometimes claimed to have saved a file without writing it, which the new check now flags.
What we are testing next
Can a slimmer bundled engine (llama.cpp) and a 2026 line-up of 2–4B models keep this reliability on an 8 GB laptop without a GPU, and can a construction 'Project Agent' pack turn a folder of site documents into a weekly report grounded in those documents?
Where it is
- Sep 2026Local chat app and auto-managed engine
- Oct 2026Teammates, routines and measured reliability layer; Windows beta
- NextSlim llama.cpp engine, small-model line-up, construction Project Agent pack
Beta. Windows only, private testers, installer not yet code-signed. Eval scores were measured on one machine (Core i9-11900KF, RTX 4070, 32 GB RAM) and are a small sample.




