The first question in an incident is how the node has been. Answering it meant reading logs. This release reads them for you and writes a report an operator or an agent can act on.
The report
- One call A scripting call returns the report; with no argument the window is the last 24 hours, with a number of hours it is that window, and zero means the whole log.
- What it contains Resource metrics, a system snapshot, disk usage, error and warning signatures aggregated, and a status computed as OK, WARNING or CRITICAL.
- Named The header names the node by host and address, resolved once and never blocking on DNS.
- Honest about coverage A coverage line states the cutoff and how many entries before it were ignored.
Reading large logs
- Skipped unopened Files last modified before the cutoff are not opened. Files over 256 MiB are read from the tail. The appender's rotation count is honoured so stray rotations are ignored.
- Robust decoding A truncated multi-byte sequence no longer fails the whole report.
The report is Markdown because its first reader is as likely to be an agent as a person.