Files
hermes-skills-devops/hermes-troubleshooting/SKILL.md
T

6.5 KiB

name, description, version, category, metadata
name description version category metadata
hermes-troubleshooting Use when Hermes gateway, auth, or platform issues arise. 1.0.0 devops
hermes
tags
hermes
troubleshooting
gateway
auth
profile
feishu

Hermes Agent Troubleshooting

Diagnose and fix common Hermes Agent issues: provider authentication failures, gateway crashes, profile isolation, and messaging platform problems.

Profile Isolation — Always Check the Right Profile

Hermes profiles have separate configs, logs, and credentials. A critical first step:

hermes profile list    # See all profiles and which gateway is running
hermes status          # Check current profile's config

Key insight: When a messaging platform (Feishu, Telegram, etc.) reports an error, the gateway is running under a different profile than the TUI session you're currently in. The TUI session's logs won't contain the platform's errors.

Log Locations Per Profile

Profile Logs Directory
default /opt/data/logs/
fey /opt/data/profiles/fey/logs/
other /opt/data/profiles/<name>/logs/

Each profile has agent.log, errors.log, and gateway.log (if gateway is running).

Provider Authentication Failed

Symptom: User on a messaging platform sees "Provider authentication failed. Check the configured credentials; raw provider details are in the gateway logs."

Root cause: This almost always means the model API key (in config.yamlmodel.api_key) is invalid/expired, NOT the platform credentials (App ID, App Secret, etc.).

Diagnostic Steps

  1. Find which profile the gateway runs under:

    hermes profile list
    # Look for the profile with "running" gateway
    
  2. Check that profile's agent.log for the real error:

    grep -i "401\|auth.*fail\|invalid.*key\|user_key" /opt/data/logs/agent.log | tail -20
    

    (Adjust path for non-default profiles.)

  3. Check the model config:

    grep "api_key\|base_url\|model:" /opt/data/config.yaml | head -10
    
  4. Check platform config (to confirm platform creds are fine):

    grep -i "feishu\|telegram\|discord" /opt/data/.env | head -10
    

Common Causes

  • Custom proxy endpoint key expired: If model.base_url points to a proxy (e.g., http://113.249.102.8:18080/...), the proxy's API key may have expired. The proxy returns 401: Authentication failed: invalid user_key.
  • API key rotated elsewhere: The key was changed in the provider's dashboard but not in Hermes .env/config.yaml.
  • Wrong profile's key: The key in the gateway profile doesn't match the one you think is active.
  • Key format mismatch when switching providers: Different providers use different key formats. A key valid for one provider (e.g., sk-sp-... for a custom proxy) will not work for another (e.g., DashScope expects sk-ant-..., DeepSeek expects sk-...). When changing model.base_url, you must also provide a key from the new provider. Verify the key with a quick curl test before restarting the gateway:
    curl -s -o /dev/null -w "%{http_code}" https://api.deepseek.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"deepseek-chat","messages":[{"role":"user","content":"hi"}],"max_tokens":1}'
    # Expect 200, not 401
    

Fix

# Update the API key in the correct profile's config
hermes -p <profile_name> config set model.api_key 'new-key'

# If also changing the provider endpoint:
hermes -p <profile_name> config set model.base_url 'https://api.provider.com/v1'
hermes -p <profile_name> config set model.default 'model-name'

# Restart gateway (may need s6 kill pattern — see references/s6-gateway-restart.md)
hermes gateway restart

hermes CLI Not on PATH

In containerized setups, hermes may not be on the default PATH. Use the full path:

/opt/hermes/.venv/bin/hermes -p default config set model.api_key 'new-key'

WebSocket 1011 and "session not found"

These are symptoms, not root causes. When the model API returns 401 repeatedly:

  • Gateway's conversation loop fails → internal error → WebSocket close code 1011
  • Session cache expires during the outage → "session not found"
  • Gateway may disconnect/reconnect the platform adapter

Fix the auth issue first; these resolve automatically.

Gateway Not Running

Symptom: hermes gateway status shows "stopped" or messaging platforms show "not configured".

Check

hermes gateway status
hermes profile list    # See which profiles have running gateways

Fix

hermes gateway run     # Foreground (for testing)
hermes gateway install # Install as background service
hermes gateway start   # Start the service

If gateway is installed but stopped (e.g., in a container with s6):

# Check if it's a container environment
hermes status | grep "Manager"
# If s6: the service may need manual start or container restart

Cannot restart from inside gateway? See references/s6-gateway-restart.md for the kill + s6 auto-restart pattern.

Gateway Crash Loop

Symptom: Gateway keeps dying and restarting.

Check crash logs

# TUI gateway crash log (thread stacks at crash)
tail -100 ~/.hermes/logs/tui_gateway_crash.log

# Systemd service failures
systemctl --user status hermes-gateway
journalctl --user -u hermes-gateway --since "1 hour ago"

Common causes

  • SSH logout kills gateway: Enable linger: sudo loginctl enable-linger $USER
  • WSL2 close kills gateway: Requires systemd=true in /etc/wsl.conf
  • Reset failed state: systemctl --user reset-failed hermes-gateway

Feishu-Specific Issues

Feishu shows "not configured" in one profile but works in another

Feishu credentials are per-profile. Check the profile that runs the gateway:

# In the gateway's profile:
grep "FEISHU" ~/.hermes/.env
# Should have: FEISHU_APP_ID, FEISHU_APP_SECRET, FEISHU_CONNECTION_MODE

Feishu WebSocket disconnects

Normal disconnects (code 1000 "bye") happen during gateway restart. Abnormal disconnects (code 1011) indicate internal errors — usually model auth failures.

Quick Diagnostic Checklist

When something is broken:

  1. hermes profile list — Which profile is running?
  2. hermes status — Is gateway running? Are platforms configured?
  3. hermes doctor — Detailed diagnostics
  4. Check the right profile's logs — not the TUI session's logs
  5. grep "401\|auth\|fail" agent.log — Find the real error behind generic messages