smartData Protection & Archiving

Upstream and subsurface

A survey shot twenty years ago may be reprocessed next quarter, so the archive has to give it back in a usable form.

Aban Smart builds archive tiers for subsurface data where the working unit is enormous, the retention horizon outlasts several hardware generations, and old data returns to active use.

Every volume figure on this page is an illustrative planning number. Survey sizes depend on acquisition geometry and vary by orders of magnitude between projects.

Very little upstream data is written once and forgotten. A field survey acquired for one purpose gets reprocessed years later with algorithms that did not exist at acquisition, and the value of that reprocessing depends entirely on whether the raw field records survived in a form that a modern processing centre can actually load. Archives built to store data rather than to return it fail this test quietly. Nobody notices until a subsurface team asks for a 2003 survey and the answer takes four months.

Industry data flow, ingest to recallData sources Seismic acquisition, Well logs, Production history ingest at Very large files. A retention driver, Asset lifecycle, decades, sets the policy that places data across Processing cache, Bulk archive, Retained master. The expected recall outcome is Re-ingest for reprocessing, hours acceptable. Tiers shown in green are governed, protected states.SOURCESINGEST & POLICYTIERSRECALLSeismicacquisitionWell logsProductionhistoryVery large filesDRIVERAsset lifecycle,decadesProcessing cacheActive reprocessingBulk archiveTape / objectRetained masterImmutableOUTCOMERe-ingest forreprocessing,hoursacceptable
Conceptual industry data flow: Seismic acquisition, Well logs, Production history ingest at Very large files, retained under Asset lifecycle, decades, placed across Processing cache, Bulk archive, Retained master, with the expected outcome Re-ingest for reprocessing, hours acceptable. Figures shown are illustrative examples, not verified customer measurements. Final topology depends on verified product compatibility.

Seismic and subsurface data volumes

Upstream storage is dominated by a small number of very large objects, which is the inverse of the financial or records estates handled elsewhere on this site. A single 3D marine survey can exceed the total annual data production of an entire corporate function, and the ratio between what was acquired and what is routinely used is unintuitive.

Data class (illustrative) Relative volume Access pattern
Raw field records Largest by far Rarely read, essential for reprocessing
Intermediate processing products Large Read during a project, then dormant
Final migrated volumes Moderate Read repeatedly by interpreters
Well logs and cores Small Read across the whole asset life
Production and reservoir history Small, continuous Appended daily for decades

The design consequence is that the small classes drive access design and the large classes drive capacity. Well logs are a rounding error in terabytes and are consulted constantly. Raw field records can be the large majority of the archive and touched once a decade. Applying one tier and one service level to both is how upstream archives become simultaneously expensive and slow.

What a multi-survey archive accumulates: a worked illustration

The numbers below are constructed to show the arithmetic. Actual survey volumes depend on fold, bin size, record length and acquisition method.

Take a single 3D survey, as an illustration:

  • Raw field records: 180 TB
  • Intermediate processing products: 45 TB
  • Final migrated volumes and deliverables: 12 TB
  • Total for one survey: 180 + 45 + 12 = 237 TB

Note the ratio. The deliverable that interpreters actually open is 12 TB, roughly 5 per cent of what has to be retained if reprocessing is to remain possible. Discarding the 180 TB looks attractive at every budget review and forecloses the option permanently.

Now extend it across an asset lifetime. Assume six surveys over thirty years, plus reprocessing campaigns:

  • Six surveys × 237 TB = 1,422 TB
  • Three reprocessing campaigns, each producing a new intermediate and final set of 57 TB: 3 × 57 = 171 TB
  • Well logs, production history and documentation over thirty years: approximately 8 TB
  • Running total: 1,422 + 171 + 8 = 1,601 TB, roughly 1.6 PB

Then the part that gets forgotten. Across thirty years the archive will be migrated between media generations perhaps four times. If each migration reads and rewrites the full archive, that is over 6 PB of verified data movement across the asset life, and each pass has a duration that must fit inside the working life of the media it is reading. A migration that takes eighteen months to move data off a format whose drives are already out of support is not a migration, it is a race. This is the core of archive migration and modernisation.

Reprocessing means re-ingestible, not just readable

There is a difference between an archive you can read and an archive you can put back to work. Restoring a 2005 survey is only useful if the processing centre can load it, which requires more than intact bytes.

Three things have to survive alongside the data:

  1. Format and header integrity. Industry seismic formats are documented, but header usage was often site-specific. Trace headers with locally defined byte locations are meaningless without the accompanying convention.
  2. Survey geometry and navigation. Without it, the traces are numbers with no position. This is the single most common reason an old survey cannot be reprocessed.
  3. Processing history. What was applied to reach the state on the media. Reprocessing decisions depend on knowing what has already been done.

Archive design should therefore treat the metadata package as a first-class object with the same retention as the data, held in an open format, and readable without the application that created it. An archive catalogue that describes each survey independently of the storage holding it is what makes a restore into a project rather than an investigation.

Storage media trade-offs: disk, tape and optical WORMFive independent scales comparing disk, LTO tape and optical write-once media on cost per terabyte, access latency, media lifespan, offline portability and idle power. Disk leads on latency only, tape on cost and power, optical on lifespan — optical is the most expensive per terabyte — so no single medium leads on every scale.MEDIA TRADE-OFFSDiskLTO tapeOptical WORMright = more favourableCost per TBacquired capacityhigherlowerAccess latencyfirst byteslowerfasterMedia lifespanbefore refreshshorterlongerPortable / offlineair-gap capablelimitedstrongPower at restidle drawhighernear zeroNo medium leads on every scale — the mix is the design decision. Positions are indicative.
Conceptual comparison. Relative positions are indicative of general media characteristics, not measured benchmarks; final selection depends on verified product specifications.

Retaining data across asset lifecycles

Corporate retention schedules usually think in years. Upstream assets think in phases: exploration, appraisal, development, plateau production, decline, then decommissioning and post-closure monitoring. Data acquired in the first phase is still consulted in the last, and some categories outlive the asset entirely.

Well integrity records, abandonment details and subsurface characterisation retain value after production stops, because responsibility for the wellbore does not end with the licence. A long-horizon subsurface retention obligation, its exact duration and its treatment at licence relinquishment, is a legal and licensing question for your own advisers, not something a storage supplier should assert.

What that means technically is that the archive must be planned for handover as well as for use. Ownership of a producing asset changes. A field is sold, a licence is relinquished, a partner exits. Each event triggers a data transfer, and the practical question is whether you can export a defined subset of the archive with its metadata and integrity manifests intact. If the answer requires reconstructing which files belong to which block, that work will be done under transaction deadline pressure.

Partner and joint-venture data obligations

Most upstream assets are operated jointly, and the data has more than one interested party. That creates requirements a single-owner archive never encounters.

  • Partition by asset, not by department. Non-operated interests and operated blocks need to be separable without a manual sift.
  • Provable non-tampering. Where partners are entitled to data supporting a cost or reserves position, checksums and write evidence answer questions that assurances do not.
  • Export packages, not access grants. Partner delivery is usually a self-describing package with a manifest, not a login to your archive.
  • Confidentiality windows. Some data becomes shareable or releasable only after a defined period, which is a retention attribute in its own right.

Legacy tape from earlier decades

Every long-lived upstream estate has a shelf of media from formats whose drives are no longer manufactured. Nine-track reels, early cartridge formats, and generations of LTO that current drives will not read. Read compatibility is limited to a small number of generations back, so an archive left untouched for long enough becomes unreadable through simple obsolescence rather than through any media failure.

Recovery is possible more often than teams expect, but it is specialist work with a real failure rate, and the honest planning assumption is that some percentage of very old media will not be recoverable. The way to avoid the problem is a scheduled refresh cycle with verification, described under tape and LTO systems. Aban Smart designs, supplies, integrates and supports archive platforms from QStar, INCOM StorEasy and DISC; we do not operate a media recovery laboratory.

Adjacent estates with similar object sizes include research and HPC and media and video. See the industries overview, or request an assessment with an inventory of your oldest media formats.

Frequently asked questions

That is a commercial judgement, but discarding raw records is irreversible and removes the option to reprocess with future algorithms. The usual compromise is to hold raw data on the cheapest durable tier with a slow retrieval service level, and keep final volumes on faster storage. Deleting the intermediate products is normally the safer saving, since they can be regenerated from raw data.

Turn your requirement into a defensible architecture

Share the workload, capacity, retention, access, and resilience requirements. Aban Smart will identify the next discovery inputs and the appropriate engagement path.