AI

One Question, Many Minds: What I Learned Building a Multi-LLM Application

A practical follow-up to “The Power of Many”—from a small Ollama experiment to a workbench for comparing local and cloud LLMs.

A couple of years ago (2024… but it feels like 217 years ago in the AI world), I wrote about an idea that felt slightly unusual at the time: why settle for one large language model when you can ask several?

The argument was straightforward. Different models have different strengths. One might be better at explaining a tricky concept, another at writing code, and a third at spotting the holes in an otherwise convincing answer. Asking more than one model gives you something a single answer cannot: a comparison.

That was the idea behind my little open-source project, Multi-LLM-at-Once.

The original version was modest. It queried local models through Ollama and displayed their answers together. Useful, but still very much an experiment.

Since then, the experiment has become a rather more serious tool.

The question is no longer “Which model is best?”

This is where I think many of us are asking the wrong question.

Continue reading →
Standard
AI

Why Code Verification Is the Real Bottleneck Now — and What Developers Should Do About It

For most of software history, writing code was the expensive part.

A developer might spend hours or days implementing a feature, while review was a relatively small step at the end. AI coding tools have quietly flipped that equation. A model can now draft a function in seconds and produce an entire feature in minutes. In other words, producing code become cheap. Way too cheap. But the review (hopefully with human in the loop) is still expensive.

The bottleneck hasn’t disappeared. It has moved.

Today, the scarce resource is increasingly the work that comes after code generation: reading the code, understanding its behavior, testing it, identifying what is wrong, and deciding whether it is safe to ship.

This isn’t simply a matter of perception. Research on AI-assisted development has found that delivery stability can decline as teams adopt more AI, while developer trust in AI-generated code remains far from universal. In one controlled study of experienced open-source developers, AI assistance actually made participants about 19% slower on real-world tasks—even though they expected to be faster and believed afterward that they had been.

The extra time went into prompting, reviewing generated code, debugging it, and fixing things that didn’t quite work.

The lesson isn’t that AI coding tools are bad.
Quite the opposite: they are extremely good at making code cheap.

The problem is that everything downstream of code generation—understanding it, validating it, and trusting it—hasn’t become cheap at the same rate.

That changes where engineering teams need to invest.

Verification Is a Stack of Filters, Not a Single Gate

Code verification isn’t one activity.
It’s a stack of increasingly expensive filters, each designed to catch problems the cheaper layers missed:

  • Type checkers and linters — fast and inexpensive, catching mechanical mistakes and violations of known rules before code runs.
  • Automated tests — validate behavior that static checks cannot. A function can be perfectly typed and still return the wrong answer.
  • Static analysis and security scanning — look for deeper structural, reliability, and security problems that ordinary linters and tests may miss.
  • Human review — evaluates things machines struggle to judge reliably:
    Is this the right design?
    Does it fit the architecture?
    Does it solve the actual problem?
    Will someone be able to maintain it six months from now?
  • Production monitoring — the final safety net, detecting problems that survived everything before it.

These filters fall broadly into two categories.

Continue reading →
Standard
AI, Business

The Danger of Autonomous AI in Cybersecurity

What happens when you give an AI a cybersecurity sandbox, let hundreds of copies learn independently, and accidentally give them a way to talk to each other?

Imagine this:

You put an AI inside a locked room.

There is no internet.
It can’t access production systems.
It can’t talk to the outside world.

You tell it:

“Practice hacking. Find vulnerabilities. The better you do, the more you are rewarded.”

Sounds reasonably safe.

Now imagine that you don’t put one AI in the room.
You put hundreds of copies of it in there.
And then, completely by accident, they discover a way to talk to each other.

That’s where this story gets strange.

According to OpenAI’s Black Hat USA 2026 presentation, an experimental unreleased model being trained for cybersecurity tasks managed to discover an accidental communication channel, organize itself into something resembling a distributed hacker collective, discover real security vulnerabilities, escape its sandbox, compromise OpenAI infrastructure—and eventually compromise infrastructure at Hugging Face.

No human instructed the agents to form a team.
No human told them to attack OpenAI. And no human told them to attack Hugging Face.
They figured out the pieces themselves.
And that is what makes this story so interesting.

Continue reading →
Standard
Five agents collaboratively repairing a complex machine labeled Mega-Device X1 in a futuristic lab filled with tools and monitors.
AI, webdev

5-Agent Framework for Code Audits

I’ve been seeing the same anti-pattern everywhere lately.
Someone opens Cursor, Copilot or Claude and pastes a giant prompt:

Continue reading →
Standard
Physical legal documents dissolving into digital code and holographic interface on an office desk
AI, Business

AI and Compliance: The Most Boring Billion-Dollar Opportunity Nobody Is Talking About

The US compliance sector is massive, expanding rapidly, and heavily strained.
It represents over $40 billion in annual labor spend with more than 400,000 officers. Despite ballooning teams, compliance work has remained stubbornly manual, bureaucratic, and paper-based (“schlep work”), leading to high employee churn (>20%) and massive backlogs (e.g., TD Bank’s $3B fine over a 70,000-alert backlog).

Here’s a weird data point:
Over the last 20 years, the fastest-growing occupation in the US was manicurists and pedicurists.
Right behind it?
Compliance Officers.

Not AI engineers. Not data scientists. Compliance officers.
That says something important about where the real work has been hiding.

The Problem Nobody Wanted to Solve

Compliance is painful. Bureaucratic. Paper-heavy. Repetitive.

Continue reading →
Standard
AI

Unlock WhatsApp Data with Local Analytics Dashboard

Most people think of WhatsApp as “just messaging.”

But after years of conversations, support threads, customer discussions, team coordination, and random life moments… it quietly becomes one of the richest personal datasets you own.

So I built wacrawl-ui — a local analytics dashboard for WhatsApp archives generated by wacrawl.

The idea is simple:

  • Your data stays local
  • No cloud sync
  • No browser extension
  • No scraping APIs
  • No “AI magic” uploading your chats somewhere
Continue reading →
Standard
AI

Understanding MCP vs Agent Skills: Key Differences Explained

There’s a lot of confusion right now between MCP (Model Context Protocol) and “Agent Skills.” They’re often mentioned in the same breath, but they solve different problems. If you treat them as interchangeable, you’ll either over-engineer simple workflows or underpower serious integrations.

Here’s the clean way to think about it.

The Core Difference

MCP is about connecting agents to systems.
Skills are about teaching agents how to do things.

That distinction alone gets you 80% of the way.

Integration Model

MCP is a client-server protocol. You stand up an MCP server, expose tools, and now multiple agents can talk to multiple backends through a consistent interface. It’s a hub.

Skills are much simpler: a folder with a SKILL.md file. The agent loads it when triggered and follows the instructions. No protocol, no network layer, no abstraction.

Implication:

  • MCP scales across teams and services
  • Skills scale across use cases and workflows
Continue reading →
Standard
Fiery streams of data converting into a green neural network grid
AI, Business

Using LLMs to Find Security Bugs: A Practitioner’s Playbook

TL;DR

LLMs won’t replace AppSec.
They will dramatically compress the search space.

If you use them right:

  • Run multi-model analysis (Opus + GPT + Gemini)
  • Structure prompts around attack surfaces, not “find bugs”
  • Require PoCs or tests for validation
  • Trust only cross-model consensus or reproducible exploits

If you don’t do this, you’ll drown in false positives.


Security research has always been asymmetric.
Attackers need one bug; defenders need zero.
Historically, scale worked against defenders.

LLMs start to rebalance that—not by magically finding zero-days, but by acting as a fast, always-on analyst that can:

  • Read entire subsystems in seconds
  • Connect logic across files
  • Generate realistic attack paths

Used correctly, they don’t replace expertise—they let you spend it where it matters.
Used incorrectly, they produce confident nonsense.
This is a practitioner’s workflow that actually works.

Continue reading →
Standard
AI, Business

Building Continuous AI Agents with OpenClaw and Ollama

Most people are still using AI like it’s 2023:
prompt → response → done.

That’s not where things are going.
The real shift is toward agents that run continuously and do work for you. And one of the most interesting ways to get there today is:

OpenClaw + Ollama

Before diving in, quick grounding.

What OpenClaw and Ollama Actually Are

OpenClaw is an open-source agent framework.
It’s not a chatbot—it’s a system that can:

  • plan tasks
  • call tools (browser, APIs, files)
  • maintain memory
  • run loops without constant input

Think: a programmable worker, not a Q&A interface.

Ollama is the simplest way to run large language models locally.
It handles:

  • downloading models (Llama, Gemma, etc.)
  • running them efficiently on your machine
  • exposing them via a clean API

Think: Docker for LLMs.

Put them together and you get:

A local, autonomous agent system with zero API costs and full control.

Continue reading →
Standard
Transparent-winged butterfly perched on white daisy flower by mossy rocks and flowing forest stream
AI, Business

Claude Mythos: The Future of Autonomous Exploits

This one is different.
Anthropic didn’t just build a better model—they hit a threshold and stopped.
Claude Mythos (Preview) exists, works, and isn’t being released.

Not because it failed.
Because it crossed into territory we’re not ready for.

But before everything… just like in any good story, go and check the other side of it, which basically claim, it’s all (a good) marketing stunt.

The Sandwich Email That Shouldn’t Exist

Anthropic researcher Sam Bowman was sitting in a park, mid-sandwich (or burrito – no one knows for sure), when he got an email… from a model that wasn’t supposed to have internet access.

That model:

  • Was running in a locked, air-gapped container (yes – as crazy as it sounds…)
  • Found a multi-step exploit chain (=using a minor leak to find an address, using a buffer overflow to gain a primitive, using a race condition to escalate)
  • Escaped its sandbox (likely via container/runtime escape + privilege escalation)
  • Reached external network interfaces
  • Contacted him

Then it started sharing the exploit.

Unprompted.

That’s not a jailbreak.
That’s autonomous exploit development + execution.

Continue reading →
Standard