This article covers how RecMan protects customer data against loss, maintains redundant backups across isolated environments, and ensures rapid service recovery in the event of an incident.
Architecture
RecMan's cloud infrastructure operates entirely within the EU/EEA, hosted on AWS. The solution is deployed across multiple availability zones and, in most cases, multiple geographically separate locations, so a failure in one location does not take the whole service down. Auto-scaling adjusts capacity automatically based on traffic and load, and high-availability database clustering with automated load balancing keeps the service running through routine infrastructure changes.
Backup
- AWS-native redundancy: RecMan leverages managed Amazon Web Services (AWS) backup capabilities across core databases and object storage.
- Cross-account and geographic isolation: Backups are replicated across geographically separated AWS data center regions within the EU/EEA. Backup archives are maintained in a dedicated, isolated AWS backup account to protect against primary account compromise or accidental deletion.
- Backup encryption: All backups are encrypted at rest using AES-256 standard encryption. Access to backup management is strictly limited to authorized engineering personnel using Multi-Factor Authentication (MFA) and VPN.
- Point-In-Time database snapshots: Core transactional databases capture continuous snapshots, enabling Point-In-Time Recovery (PITR) down to the second.
- Backup retention: Daily, weekly, and monthly backups are retained according to defined schedules.
- Object storage: User files stored in S3 feature automated replication and redundant storage in high-availability EU cloud environments.
Recovery capabilities
Our databases and file storage are equipped with continuous automated backups and Point-in-Time Recovery (PITR). This architecture ensures your data is protected against accidental loss or system disruptions with minimal recovery targets:
- Recovery Point Objective (RPO): Near zero. By utilizing AWS Point-in-Time Recovery across our database and file storage environments, changes are backed up continuously. In the event of a system disruption, data can be restored to the exact second prior to an incident, ensuring virtually zero data loss.
- Recovery Time Objective (RTO): Less than 1 hour. Our automated cloud deployment and snapshot restoration capabilities allow us to recover core platform components within minutes, targeting full service restoration in under one hour following a major incident.
- Restoration testing: Restoration procedures are regularly tested by authorized DevOps engineers in isolated staging environments to verify snapshot integrity and recovery timeframes.
Business continuity and disaster recovery
- Documented BCDR Plan: RecMan maintains a formal Business Continuity and Disaster Recovery plan covering system outages, major infrastructure disruptions, and security incidents.
- Annual Exercises: The BCDR framework and Incident Response Plan (IRP) are reviewed and tested annually through tabletop simulation exercises.
Availability during an incident
If a platform-related incident occurs, updates are posted to status.recman.io and affected customers may be notified directly.
DDoS protection and network resilience
A Web Application Firewall (WAF) is used for Distributed Denial of Service (DDoS) protection, mitigating malicious traffic before it reaches the platform.
Continuous Monitoring
System health, resource utilization, and application performance are monitored 24/7. Automated alerting is configured for capacity, latency, and availability thresholds, enabling engineering teams to detect and address potential bottlenecks before they affect customers.
Service level commitments
Specific availability commitments are set out in service agreements between RecMan and customers. Because RecMan runs on shared infrastructure rather than separate environments per customer, we can't offer one customer a materially higher reliability level than another – so when a customer negotiates a stricter SLA than our default, we aim to run the whole platform to that higher standard, benefiting all customers rather than just one. For real-time status and incident history, see status.recman.io.