smartData Protection & Archiving

Readable in fifteen years

An archive is not storage you stop thinking about. It is storage with a maintenance schedule.

Data outlives the media, the drives, the software and usually the team. Designing for that horizon means planning refresh cycles and recall times up front, not discovering them during an audit.

Aban Smart designs, supplies, integrates and supports archive tiers built on QStar software with INCOM/StorEasy and DISC hardware across the GCC.

Ask when an archive last had a file read out of it and the honest answer is often that nobody knows. That is what makes long-horizon storage awkward. It fails quietly, and the failure surfaces in front of whoever needs the record most urgently. Fifteen years spans four or five LTO generations, several software major versions and at least one full turnover of the people who built the thing. A design that survives it is organised around scheduled change rather than around capacity, and the cost of that change belongs in the budget from year one.

Data lifecycle timelineA horizontal timeline running ingest, active, nearline, archive and expiry, with a policy gate between each stage for classification, age policy, retention and legal hold. Access frequency falls from constant to none across the top; the tier each stage lands on runs along the bottom, from flash and disk to tape and optical WORM. Expiry is a defensible disposition with a certificate of erasure and a retained audit log.ACCESS FREQLANDS ONCLASSIFYAGE POLICYRETENTIONLEGAL HOLDConstantIngestcaptureFlash / diskDailyActivein productionDisk / objectWeeklyNearlineoccasional useObject / tapeRareArchiveretainedTape + opticalWORMNoneExpirydispositionCertified eraseDefensible disposition — not just deletioncertificate of erasure, audit log retained, legal hold honouredConceptual lifecycle. Gate criteria and retention periods depend on policy and regulation.
Conceptual data lifecycle. Gate criteria and retention periods depend on applicable policy and regulation.

Policy gates decide when data moves, not people.

Passive shelf versus active archive

A passive archive is a write-and-hope arrangement. Files are copied to media, the media goes somewhere safe, and an index lives in a spreadsheet maintained by whoever was on shift. Retrieval is a request, a person, a search and often a day. It is cheap until the moment it is needed.

An active archive keeps the data offline or nearline but keeps knowledge of it online. A catalogue holds what exists, how many copies were written, which volume each copy sits on and whether a checksum still matches. Applications see one path that does not change when the medium underneath changes. A read against that path blocks for a while and then returns the file, without the requester learning that a robot moved anything.

The distinction is not marketing. The Active Archive Alliance, which QStar rejoined in February 2020 having been one of four founding members in 2010, frames it as online access to data across every tier of the hierarchy. The practical version is simpler. If finding something requires a human to remember where it went, you have a shelf. If the system knows, you have an archive. The mechanics of the catalogue and the stub layer are described under archive management software.

Format obsolescence and media refresh

Bit rot gets the attention. It is rarely what strands an archive. The failures that actually happen are structural and entirely predictable.

  • Drive generation compatibility. A current LTO drive reads only one or two generations of older media, and new generations arrive every two to three years. Media that is physically perfect becomes unreadable because no supported drive in the building can mount it.
  • Proprietary containers. An optimised vendor format may pack more onto a cartridge than an open one. It also means the only route to your data is a product licence that must stay current for the life of the retention period. LTFS exists as the hedge against that, and choosing between the two is a decision that is very hard to revisit after several thousand cartridges.
  • Catalogue loss. Media without its index is a set of opaque volumes. The catalogue needs its own protected copy and its own restore rehearsal, separately from the data it describes.
  • Software and supplier discontinuity. End of life, acquisition, or a product line quietly dropped. The question to ask a supplier is not whether this will happen but what the export path looks like when it does.

None of that is fixed by buying better media. It is fixed by treating refresh as a recurring operation with a date and a budget line, driven by drive support horizons rather than by the rated shelf life printed on a box. Because the read window is narrow and the release cadence is not, a useful planning habit is to budget for copying the whole set forward every five to eight years, and to price that before the first cartridge is bought.

Refresh also consolidates. Because each generation holds more, a many-to-one copy usually reduces the number of physical items you look after, which lowers handling cost for the cycle after. Where the source is old enough that readability is in doubt, the work becomes a project in its own right, covered under archive migration and modernisation.

Recall time as a design input

Retrieval speed is an architectural requirement, not a footnote. It should be stated per data class before hardware is selected, because it is the input that decides how much disk cache you buy and whether removable media is viable at all.

The table below is illustrative. It describes typical first-byte behaviour by tier, not measured figures from any specific installation.

Tier What happens before the first byte Typical order of magnitude Dominant variable
Flash or disk cache Nothing. The copy is already local Milliseconds Cache sizing and hit rate
On-site object storage Network round trip and object lookup Sub-second to seconds Node load and network path
Online tape in a library Catalogue lookup, robotic pick, load, thread, seek Tens of seconds to a few minutes Free drive availability
Optical library Catalogue lookup, robotic pick, spin-up, seek Tens of seconds to minutes Drive count in the library
Offline media on a shelf A person locates, transports and loads the item Hours to days Custody procedure and staffing
Cloud archive storage class Provider restore job, then transfer Minutes to hours before transfer begins Provider tier and retrieval option chosen

The column that matters most in practice is the last one. A single tape mount is fast enough for almost anybody. Forty simultaneous requests against four drives is a queue, and the fortieth user waits for thirty-nine mounts ahead of them. Concurrency and drive count set the experience, not the mount time quoted in a datasheet.

Human handling dominates at the bottom of the table by an even wider margin. An offline copy retrieved from a vault has a recall time measured by transport and authorisation, not by hardware. That is acceptable for a terminal retention copy and unacceptable for anything a clinician or a case worker requests during a working day. Tier by request pattern, not by age alone.

Storage media trade-offs: disk, tape and optical WORMFive independent scales comparing disk, LTO tape and optical write-once media on cost per terabyte, access latency, media lifespan, offline portability and idle power. Disk leads on latency only, tape on cost and power, optical on lifespan — optical is the most expensive per terabyte — so no single medium leads on every scale.MEDIA TRADE-OFFSDiskLTO tapeOptical WORMright = more favourableCost per TBacquired capacityhigherlowerAccess latencyfirst byteslowerfasterMedia lifespanbefore refreshshorterlongerPortable / offlineair-gap capablelimitedstrongPower at restidle drawhighernear zeroNo medium leads on every scale — the mix is the design decision. Positions are indicative.
Conceptual comparison. Relative positions are indicative of general media characteristics, not measured benchmarks; final selection depends on verified product specifications.

No medium leads on every axis.

Cloud as one tier with a price list

Cloud archive classes belong in the design. They are a poor default and a good component, and the difference is in the arithmetic.

Storage cost per terabyte is the number most easily compared and the least likely to dominate. Retrieval charges, request charges, minimum storage duration commitments and egress fees are what turn an inexpensive archive into an expensive recovery. A restore of a few files costs nothing worth discussing. A bulk restore of a hundred terabytes under time pressure is a budget event, and it arrives during an incident when nobody wants a procurement conversation.

Two further constraints apply before economics. Retrieval latency in the deepest classes is measured in hours, which is fine for a disaster copy and wrong for a working archive. Residency may limit which regions are permitted for your data at all. Where those hold, an on-premises tier with a cloud copy behind it is usually a better shape than the reverse.

Deciding what moves and what leaves

Lifecycle policy is the part organisations skip, and its absence is why archives grow monotonically. Classification tools work from attributes storage can see for itself: age since creation, time since last access, size, extension, owner, location in the tree. No tagging effort is required from users, which is the only reason such rules survive contact with reality.

Sensible policy has three actions, not one. Move data off primary when it stops being touched. Keep a second copy on different technology, because one archive copy is a single point of failure wearing a reassuring name. And expire data when its period ends, which requires somebody with the authority to approve disposal. Where that authority is undefined, expiry stalls and the archive becomes permanent by accident rather than by decision. Retention classes, holds and the evidence around disposal sit under regulatory data retention.

Residency, imports and the GCC operating reality

Two constraints shape archive design in the UAE and the wider Gulf more than any technical preference.

The first is location. Sector regulators and boards frequently require records to remain in-country, which narrows the cloud region list quickly and sometimes empties it. Where the requirement is firm, the archive tier is on-premises and the cloud discussion moves to a secondary copy, if it happens at all.

The second is logistics. Media, drives and library spares are imported. A refresh cycle that assumes next-day parts is planning against a supply chain that does not exist here, so lead times belong in the schedule rather than in the risk register. Environmental conditions matter too: media stored in a conditioned data hall behaves differently from media in an office store room during a Gulf summer, and the storage location for offline copies should be specified in the design rather than improvised later.

For the media layer itself, see tape and LTO systems. The wider set of related work is listed under solutions. If you want an existing archive examined for refresh exposure and recall behaviour, request an assessment.

Related resources

Manufacturer documentation relevant to this page. Availability, specifications, and configurations are subject to verification.

PDFDatasheet
Date not verified

QStar Archive Manager Datasheet

Archive management software that virtualizes tape, optical, disk, object, and cloud targets behind file and S3 interfaces. Confirm supported versions and configuration before procurement.

Vendor:
QStar
Date:
Jun 2024
Format:
PDF
Size:
949 KB

Frequently asked questions

Cold cloud storage is a destination. An active archive is a managed namespace that can place data on several destinations, including cold cloud, while keeping a single searchable index and a stable path for applications. The practical difference shows up on retrieval and on refresh: an active archive can move data to new media without changing anything an application sees.

Turn your requirement into a defensible architecture

Share the workload, capacity, retention, access, and resilience requirements. Aban Smart will identify the next discovery inputs and the appropriate engagement path.