Introduction to advanced administration
Table of contents
- The senior administrator role
- Essential skills
- Production environments
- Methodology and best practices
- Administrator tools
- Practical exercises
1 - The senior administrator role
Responsibilities
Differences from a junior admin
| Aspect | Junior | Senior |
|---|---|---|
| Scope | Assigned tasks | Big-picture view |
| Problems | Guided resolution | Autonomous diagnosis |
| Decisions | Follows procedures | Defines procedures |
| Architecture | Implements | Designs |
| Security | Applies the rules | Defines the policy |
Career progression
🔝 Back to table of contents
2 - Essential skills
Hard Skills
| Domain | Skills |
|---|---|
| System | Kernel, systemd, performance tuning |
| Network | Advanced TCP/IP, firewall, VPN, load balancing |
| Storage | RAID, LVM, SAN, NFS, iSCSI |
| Security | Hardening, audit, SELinux/AppArmor |
| Virtualization | KVM, VMware, containers |
| Automation | Bash, Python, Ansible, Terraform |
| Monitoring | Prometheus, Grafana, ELK |
Soft Skills
- Communication: Explaining technical concepts
- Documentation: Clear and maintained procedures
- Stress management: Production incidents
- Mentoring: Training juniors
- Strategic vision: Anticipating needs
Recognized certifications
| Certification | Vendor | Level |
|---|---|---|
| RHCSA | Red Hat | Intermediate |
| RHCE | Red Hat | Advanced |
| LFCS | Linux Foundation | Intermediate |
| LFCE | Linux Foundation | Advanced |
| CompTIA Linux+ | CompTIA | Intermediate |
🔝 Back to table of contents
3 - Production environments
Characteristics of a production environment
SLA levels
| SLA | Downtime/year | Downtime/month |
|---|---|---|
| 99% | 3.65 days | 7.3 hours |
| 99.9% | 8.76 hours | 43.8 minutes |
| 99.99% | 52.6 minutes | 4.38 minutes |
| 99.999% | 5.26 minutes | 26.3 seconds |
Important
A 99.9% SLA sounds high, but it still represents 43 minutes of downtime per month!
Typical environments
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ DEV │ → │ STAGING │ → │ PROD │
│ │ │ │ │ │
│ Développeurs│ │ Tests QA │ │ Utilisateurs│
│ Libre accès │ │ Pré-prod │ │ Accès limité│
└─────────────┘ └─────────────┘ └─────────────┘
🔝 Back to table of contents
4 - Methodology and best practices
ITIL - Service management
| Process | Description |
|---|---|
| Incident Management | Restore the service quickly |
| Problem Management | Identify and eliminate root causes |
| Change Management | Control changes |
| Release Management | Deploy reliably |
Change management
Mandatory documentation
# Structure de documentation recommandée
docs/
├── architecture/
│ ├── infrastructure.md
│ └── network-diagram.md
├── procedures/
│ ├── backup-restore.md
│ ├── incident-response.md
│ └── deployment.md
├── runbooks/
│ ├── service-restart.md
│ └── troubleshooting-guide.md
└── policies/
├── security-policy.md
└── change-management.md
Post-mortem after an incident
# Post-mortem : [Titre de l'incident]
## Résumé
- Date : YYYY-MM-DD
- Durée : X heures
- Impact : [Description]
## Timeline
- HH:MM - Détection
- HH:MM - Actions prises
- HH:MM - Résolution
## Cause racine
[Explication détaillée]
## Actions correctives
1. [ ] Action 1
2. [ ] Action 2
## Leçons apprises
[Ce qu'on retient]
🔝 Back to table of contents
5 - Administrator tools
Monitoring and observability
| Tool | Usage |
|---|---|
| Prometheus | Metrics and alerts |
| Grafana | Dashboards |
| ELK Stack | Centralized logs |
| Nagios/Zabbix | Infrastructure monitoring |
| Datadog/New Relic | SaaS APM |
Configuration Management
| Tool | Type |
|---|---|
| Ansible | Agentless, YAML |
| Puppet | Agent, Ruby DSL |
| Chef | Agent, Ruby DSL |
| SaltStack | Agent/Agentless, YAML |
Infrastructure as Code
| Tool | Usage |
|---|---|
| Terraform | Multi-cloud provisioning |
| CloudFormation | AWS IaC |
| Pulumi | IaC with general-purpose languages |
Essential CLI tools
# Performance
htop, iotop, iftop, nethogs, dstat
# Réseau
tcpdump, wireshark, nmap, ss, ip
# Système
strace, ltrace, lsof, fuser
# Logs
journalctl, tail, less, grep, awk
# Stockage
lsblk, fdisk, parted, pvs, vgs, lvs
🔝 Back to table of contents
6 - Practical exercises
Exercise 1: Self-assessment
Assess your current skills:
| Domain | 1-5 | To improve |
|---|---|---|
| Bash scripting | ||
| Network configuration | ||
| System security | ||
| Monitoring | ||
| Troubleshooting |
Exercise 2: Documentation
Create a runbook for restarting a critical service:
Suggested structure
# Runbook : Redémarrage du service [NOM]
## Prérequis
- Accès root au serveur
- Vérifier la fenêtre de maintenance
## Procédure
1. Vérifier l'état actuel
2. Notifier les équipes
3. Arrêter le service
4. Vérifier les logs
5. Démarrer le service
6. Valider le fonctionnement
7. Clôturer
## Rollback
En cas de problème...
Quiz
Q1. What does a 99.99% SLA mean in terms of monthly downtime?
Answer
About 4.38 minutes of downtime per month maximum.
Q2. What is the difference between Incident Management and Problem Management?
Answer
- Incident Management: Restore the service quickly (reactive)
- Problem Management: Identify and eliminate the root cause (proactive)
🔝 Back to table of contents
Key takeaways
- The senior admin has a big-picture view and defines procedures
- SLAs dictate the availability requirements
- Documentation is as important as technical skills
- Change Management prevents production incidents
- Master monitoring and automation tools