With Zabbix now taking over from Nagios, we need to document it better. I'm happy to deal with that. I think the initial list would be:
SOPs - Taking a single machine in/out of monitoring for reboots etc - Scheduling a maintenance window for a group of hosts - Updating & applying template variables (eg cpu load) for a host or group via Ansible
Larger Howtos - Overview of Zabbix and how we've implemented it - Developing a new template for a group of checks on certain hosts - SAML mappings and which groups have what levels of access
Any others we need?
Metadata Update from @gwmngilfen: - Issue assigned to gwmngilfen
Metadata Update from @gwmngilfen: - Issue tagged with: monitoring
Just wanted to let you know that the documentation is now live in https://forge.fedoraproject.org/infra/docs
Metadata Update from @james: - Issue priority set to: Waiting on Assignee (was: Needs Review) - Issue tagged with: low-gain, medium-trouble
Metadata Update from @gwmngilfen: - Issue tagged with: sprint-0
This issue has been migrated to Fedora Forge: https://forge.fedoraproject.org/infra/tickets/issues/12977
Please continue any further discussion there.
Metadata Update from @ryanlerch: - Issue close_status updated to: Migrated to Fedora Forge - Issue status updated to: Closed (was: Open)