My Hermes Agents Looked Healthy — But Couldn’t Actually Think
Everything looked alive
I run several specialised Hermes agents on a Mac mini in my homelab.
Each one has a deliberately limited job. Infrastructure Sentinel looks at infrastructure health, Network Detective investigates networking, Backup & Recovery Engineer reasons about backups, Home Intelligence Engineer works with deterministic smart-home evidence, and Change Reviewer reviews proposed production changes.
They each run through their own local gateway and are exposed through a Home Assistant dashboard I built called Hermes Operations.
After updating Hermes to version 0.21.3, something strange happened.
Several of the agents stopped working in Home Assistant.
At first, the obvious infrastructure did not look broken. Their gateway processes could still be running. Ports could still be listening. From a traditional service-health perspective, parts of the system appeared alive. 08_HERMES_MULTI_AGENT_OPERATION…
But the agents themselves could no longer actually do their job.
A healthy process is not a healthy agent
When I started investigating the individual Hermes profiles, I found the important clue.
For several specialists:
hermes -p PROFILE auth listreturned no usable OpenAI/Codex authentication.
The gateways were there, the processes were there. But the profiles could not actually reach the reasoning model.
The specialist profiles had effectively become isolated authentication islands instead of inheriting the authentication state I expected them to use. I restored the affected agents by adding a supported OpenAI/Codex OAuth credential to each profile and restarting its gateway.
After that, Infrastructure Sentinel, Network Detective, Backup & Recovery Engineer and Home Intelligence Engineer worked again. Change Reviewer already had its own credential.
One important thing I deliberately did not conclude was that “the tokens periodically expire”.
I did not have evidence for that.
The incident happened around a Hermes update and profile/authentication behaviour, so it may have been a one-time migration or update-related problem. I would rather keep that uncertainty explicit than invent a cleaner root cause than the evidence supports.
The problem with /health
This incident exposed a bigger observability problem.
For a normal web service, checking whether the process is alive and whether an HTTP endpoint returns 200 OK might already tell you quite a lot. For an AI agent, that is not enough.
I now think of agent health as three separate layers:
Gateway
↓
Credential
↓
Real inferenceA gateway can be reachable while authentication is missing. A credential record can exist while the provider still rejects it.
And even if both of those look fine, the only way to prove that the agent can actually think is to make it perform a real inference request.
That led me to build Hermes Auth Health v0.1.
Testing whether the agent can actually think
Auth Health is deliberately not another AI agent. It is a deterministic Python checker. For every fixed production specialist, it performs three checks.
First, it checks the agent gateway:
Is the endpoint reachable?
Does it return valid JSON?
Does it identify itself as Hermes?Then it checks authentication:
Does this profile actually contain
an OpenAI/Codex credential?And finally, it performs a real one-shot inference request.
The model receives a fixed prompt and must return exactly:
HERMES_INFERENCE_OKOnly when all three checks pass do I consider that specialist healthy.
That final step is the important one.
I am not asking:
Is the AI server running?
I am asking:
Can this exact production agent actually perform inference right now?
Five out of five
After building and testing the checker, I ran it against the full production specialist set.
The result was:
Infrastructure Sentinel HEALTHY
Network Detective HEALTHY
Backup & Recovery Engineer HEALTHY
Home Intelligence Engineer HEALTHY
Change Reviewer HEALTHY
5/5 HEALTHYFor each one, gateway health, credential presence and real inference all passed.
The checker itself is also intentionally read-only. It does not restart gateways, repair credentials or change configuration when something fails.
Its job is to tell me what is true.
Repair is a separate decision.
Why it still does not run automatically
The slightly funny part is that Auth Health itself uncovered another issue during testing.
The inference probes do not appear in the normal Hermes session list, but they are still persisted internally.
One complete checker run creates one internal inference session per profile. During production acceptance, two runs across five profiles created ten new sessions and twenty message rows.
That is not a disaster, but I do not want to schedule a health check every few minutes and quietly fill the agent database with monitoring sessions.
So Auth Health is currently manual and on-demand. No cron job. No automatic repair. No notification pipeline. Not yet.
I would rather leave a known limitation visible than automate something before I understand its long-term behaviour.
What I learned
The interesting part of this incident was not really OAuth.
It was the definition of health.
A process being alive does not prove the service works.
A gateway being reachable does not prove the agent works.
A credential existing does not prove authentication works.
And for an AI system, even all of that does not prove much until the model actually responds.
So my health model is now:
process health ≠ authentication health ≠ inference health
That is a distinction I probably would not have thought much about before building agents into my homelab.
But once AI becomes part of real infrastructure, “the port is open” is no longer a very convincing definition of healthy.