Skip to main content

Introduction to advanced administration


Table of contents

  1. The senior administrator role
  2. Essential skills
  3. Production environments
  4. Methodology and best practices
  5. Administrator tools
  6. Practical exercises


1 - The senior administrator role

Responsibilities

Differences from a junior admin

AspectJuniorSenior
ScopeAssigned tasksBig-picture view
ProblemsGuided resolutionAutonomous diagnosis
DecisionsFollows proceduresDefines procedures
ArchitectureImplementsDesigns
SecurityApplies the rulesDefines the policy

Career progression

🔝 Back to table of contents



2 - Essential skills

Hard Skills

DomainSkills
SystemKernel, systemd, performance tuning
NetworkAdvanced TCP/IP, firewall, VPN, load balancing
StorageRAID, LVM, SAN, NFS, iSCSI
SecurityHardening, audit, SELinux/AppArmor
VirtualizationKVM, VMware, containers
AutomationBash, Python, Ansible, Terraform
MonitoringPrometheus, Grafana, ELK

Soft Skills

  • Communication: Explaining technical concepts
  • Documentation: Clear and maintained procedures
  • Stress management: Production incidents
  • Mentoring: Training juniors
  • Strategic vision: Anticipating needs

Recognized certifications

CertificationVendorLevel
RHCSARed HatIntermediate
RHCERed HatAdvanced
LFCSLinux FoundationIntermediate
LFCELinux FoundationAdvanced
CompTIA Linux+CompTIAIntermediate

🔝 Back to table of contents



3 - Production environments

Characteristics of a production environment

SLA levels

SLADowntime/yearDowntime/month
99%3.65 days7.3 hours
99.9%8.76 hours43.8 minutes
99.99%52.6 minutes4.38 minutes
99.999%5.26 minutes26.3 seconds
Important

A 99.9% SLA sounds high, but it still represents 43 minutes of downtime per month!

Typical environments

┌─────────────┐    ┌─────────────┐    ┌─────────────┐
│ DEV │ → │ STAGING │ → │ PROD │
│ │ │ │ │ │
│ Développeurs│ │ Tests QA │ │ Utilisateurs│
│ Libre accès │ │ Pré-prod │ │ Accès limité│
└─────────────┘ └─────────────┘ └─────────────┘

🔝 Back to table of contents



4 - Methodology and best practices

ITIL - Service management

ProcessDescription
Incident ManagementRestore the service quickly
Problem ManagementIdentify and eliminate root causes
Change ManagementControl changes
Release ManagementDeploy reliably

Change management

Mandatory documentation

# Structure de documentation recommandée
docs/
├── architecture/
│ ├── infrastructure.md
│ └── network-diagram.md
├── procedures/
│ ├── backup-restore.md
│ ├── incident-response.md
│ └── deployment.md
├── runbooks/
│ ├── service-restart.md
│ └── troubleshooting-guide.md
└── policies/
├── security-policy.md
└── change-management.md

Post-mortem after an incident

# Post-mortem : [Titre de l'incident]

## Résumé
- Date : YYYY-MM-DD
- Durée : X heures
- Impact : [Description]

## Timeline
- HH:MM - Détection
- HH:MM - Actions prises
- HH:MM - Résolution

## Cause racine
[Explication détaillée]

## Actions correctives
1. [ ] Action 1
2. [ ] Action 2

## Leçons apprises
[Ce qu'on retient]

🔝 Back to table of contents



5 - Administrator tools

Monitoring and observability

ToolUsage
PrometheusMetrics and alerts
GrafanaDashboards
ELK StackCentralized logs
Nagios/ZabbixInfrastructure monitoring
Datadog/New RelicSaaS APM

Configuration Management

ToolType
AnsibleAgentless, YAML
PuppetAgent, Ruby DSL
ChefAgent, Ruby DSL
SaltStackAgent/Agentless, YAML

Infrastructure as Code

ToolUsage
TerraformMulti-cloud provisioning
CloudFormationAWS IaC
PulumiIaC with general-purpose languages

Essential CLI tools

# Performance
htop, iotop, iftop, nethogs, dstat

# Réseau
tcpdump, wireshark, nmap, ss, ip

# Système
strace, ltrace, lsof, fuser

# Logs
journalctl, tail, less, grep, awk

# Stockage
lsblk, fdisk, parted, pvs, vgs, lvs

🔝 Back to table of contents



6 - Practical exercises

Exercise 1: Self-assessment

Assess your current skills:

Domain1-5To improve
Bash scripting
Network configuration
System security
Monitoring
Troubleshooting

Exercise 2: Documentation

Create a runbook for restarting a critical service:

Suggested structure
# Runbook : Redémarrage du service [NOM]

## Prérequis
- Accès root au serveur
- Vérifier la fenêtre de maintenance

## Procédure
1. Vérifier l'état actuel
2. Notifier les équipes
3. Arrêter le service
4. Vérifier les logs
5. Démarrer le service
6. Valider le fonctionnement
7. Clôturer

## Rollback
En cas de problème...

Quiz

Q1. What does a 99.99% SLA mean in terms of monthly downtime?

Answer

About 4.38 minutes of downtime per month maximum.

Q2. What is the difference between Incident Management and Problem Management?

Answer
  • Incident Management: Restore the service quickly (reactive)
  • Problem Management: Identify and eliminate the root cause (proactive)

🔝 Back to table of contents



Key takeaways

  • The senior admin has a big-picture view and defines procedures
  • SLAs dictate the availability requirements
  • Documentation is as important as technical skills
  • Change Management prevents production incidents
  • Master monitoring and automation tools

🔝 Back to table of contents


Next chapter: Hardening and security →