Articles
Practical breakdowns of engineering systems: autonomous AI teams, QA workflows, releases, infrastructure, and the constraints without which automation quickly turns into a convincing illusion.
The Perfect Task for Jev: A Classification Model Lost to General-Purpose LLMs
I have a task that reads like a sales example for Jev.
My bot, Derzhi Lida (Russian for “grab the lead”), reads more than 300 working chats in Telegram and looks for client requests and job openings for paid-ads and social media specialists. Regular expressions drop the obvious noise, and a model decides the rest. About 1,600 decisions a day. The answer is always one of three: a client request, a job opening, or irrelevant. Plus a list of ad platforms. Nothing to write, only something to choose.
Read →How I Handed My AI Tech Teams to an AI Lead
A follow-up to the July article on autonomous AI tech teams: an AI lead now sits above the teams, decides from my precedents, catches sessions that stop for good and checks every closure itself.
Read →How AI Turned Five ASICs in a Country House Into a Self-Funding Climate System
A country property contains five independent heating zones. Each has a quiet Jasminer X16-Q that converts its electrical input into heat while mining Ethereum Classic.
Climate is managed by a local AI autopilot with outdoor weather, water-pipe freeze protection, fan and hardware-switch monitoring, live heating-cost calculations, and a Telegram control panel. It runs on a local Linux server, distributes up to three “heat slots” across the zones, and does not depend on a cloud-hosted model.
Read →Choosing a Fallback LLM for Plan Review: Benchmarking 12 Models and Adopting Tencent Hunyuan 3
In autonomous agent development (Agentic SDD), one of the most critical stages is reviewing implementation plans and specs before writing any code. If a logical defect slips past the planning stage, the implementing agent wastes dozens of minutes and API dollars implementing a broken architecture, while QA agents inherit false-positive test assertions.
Our primary engine for deep plan review is the OpenAI Codex CLI running at maximum reasoning effort (reasoning_effort="xhigh"). However, in fully autonomous overnight batches, we periodically hit token quotas and provider rate limits. Furthermore, single-vendor dependency represented a single point of failure.
Agentic SDD: Forcing AI Agents to Write Code by Strict Contract
In Findrates.ai, Agentic SDD (autonomous spec-driven development) runs every day: we have closed over 240 plans through this pipeline. Our actual process differs significantly from conference slides where classical SDD is presented as an effortless silver bullet.
The Limits of Vibecoding
With the rise of LLMs, developers initially relied on “vibecoding”—prompting a chatbot to “build feature X.” For quick scripts and prototypes, this was enough. On an existing codebase, however, vibecoding quickly caused loss of context: models forgot constraints between sessions, broke neighboring modules, and generated unmanaged technical debt.
Read →18 Days Without a Single Task From Me
An owner’s report after 18 days: how an autonomous employee runs on schedule, closes its own hypotheses, catches its own mistakes, and stops where the decision belongs to a human.
Read →How I Built a Concierge for Claude Code and Codex
A practical breakdown of a personal AI concierge for Claude Code and Codex: how it understands incoming requests, resumes the right sessions, tracks assignments, and returns results.
Read →How I Built an Autonomous Digital Employee
A practical breakdown of an autonomous digital employee: schedule, authority, memory, engineering intuition, a research pipeline, and risk boundaries.
Read →How I Built a Workflow for Autonomous AI Engineering Teams
A practical breakdown of a system of autonomous AI engineering teams: roles, task queue, QA, audit, retrospectives, and the release loop.
Read →