Upstream and subsurface
A survey shot twenty years ago may be reprocessed next quarter, so the archive has to give it back in a usable form.
Aban Smart builds archive tiers for subsurface data where the working unit is enormous, the retention horizon outlasts several hardware generations, and old data returns to active use.
Every volume figure on this page is an illustrative planning number. Survey sizes depend on acquisition geometry and vary by orders of magnitude between projects.
Very little upstream data is written once and forgotten. A field survey acquired for one purpose gets reprocessed years later with algorithms that did not exist at acquisition, and the value of that reprocessing depends entirely on whether the raw field records survived in a form that a modern processing centre can actually load. Archives built to store data rather than to return it fail this test quietly. Nobody notices until a subsurface team asks for a 2003 survey and the answer takes four months.
Seismic and subsurface data volumes
Upstream storage is dominated by a small number of very large objects, which is the inverse of the financial or records estates handled elsewhere on this site. A single 3D marine survey can exceed the total annual data production of an entire corporate function, and the ratio between what was acquired and what is routinely used is unintuitive.
| Data class (illustrative) | Relative volume | Access pattern |
|---|---|---|
| Raw field records | Largest by far | Rarely read, essential for reprocessing |
| Intermediate processing products | Large | Read during a project, then dormant |
| Final migrated volumes | Moderate | Read repeatedly by interpreters |
| Well logs and cores | Small | Read across the whole asset life |
| Production and reservoir history | Small, continuous | Appended daily for decades |
The design consequence is that the small classes drive access design and the large classes drive capacity. Well logs are a rounding error in terabytes and are consulted constantly. Raw field records can be the large majority of the archive and touched once a decade. Applying one tier and one service level to both is how upstream archives become simultaneously expensive and slow.
What a multi-survey archive accumulates: a worked illustration
The numbers below are constructed to show the arithmetic. Actual survey volumes depend on fold, bin size, record length and acquisition method.
Take a single 3D survey, as an illustration:
- Raw field records: 180 TB
- Intermediate processing products: 45 TB
- Final migrated volumes and deliverables: 12 TB
- Total for one survey: 180 + 45 + 12 = 237 TB
Note the ratio. The deliverable that interpreters actually open is 12 TB, roughly 5 per cent of what has to be retained if reprocessing is to remain possible. Discarding the 180 TB looks attractive at every budget review and forecloses the option permanently.
Now extend it across an asset lifetime. Assume six surveys over thirty years, plus reprocessing campaigns:
- Six surveys × 237 TB = 1,422 TB
- Three reprocessing campaigns, each producing a new intermediate and final set of 57 TB: 3 × 57 = 171 TB
- Well logs, production history and documentation over thirty years: approximately 8 TB
- Running total: 1,422 + 171 + 8 = 1,601 TB, roughly 1.6 PB
Then the part that gets forgotten. Across thirty years the archive will be migrated between media generations perhaps four times. If each migration reads and rewrites the full archive, that is over 6 PB of verified data movement across the asset life, and each pass has a duration that must fit inside the working life of the media it is reading. A migration that takes eighteen months to move data off a format whose drives are already out of support is not a migration, it is a race. This is the core of archive migration and modernisation.
Reprocessing means re-ingestible, not just readable
There is a difference between an archive you can read and an archive you can put back to work. Restoring a 2005 survey is only useful if the processing centre can load it, which requires more than intact bytes.
Three things have to survive alongside the data:
- Format and header integrity. Industry seismic formats are documented, but header usage was often site-specific. Trace headers with locally defined byte locations are meaningless without the accompanying convention.
- Survey geometry and navigation. Without it, the traces are numbers with no position. This is the single most common reason an old survey cannot be reprocessed.
- Processing history. What was applied to reach the state on the media. Reprocessing decisions depend on knowing what has already been done.
Archive design should therefore treat the metadata package as a first-class object with the same retention as the data, held in an open format, and readable without the application that created it. An archive catalogue that describes each survey independently of the storage holding it is what makes a restore into a project rather than an investigation.
Retaining data across asset lifecycles
Corporate retention schedules usually think in years. Upstream assets think in phases: exploration, appraisal, development, plateau production, decline, then decommissioning and post-closure monitoring. Data acquired in the first phase is still consulted in the last, and some categories outlive the asset entirely.
Well integrity records, abandonment details and subsurface characterisation retain value after production stops, because responsibility for the wellbore does not end with the licence. A long-horizon subsurface retention obligation, its exact duration and its treatment at licence relinquishment, is a legal and licensing question for your own advisers, not something a storage supplier should assert.
What that means technically is that the archive must be planned for handover as well as for use. Ownership of a producing asset changes. A field is sold, a licence is relinquished, a partner exits. Each event triggers a data transfer, and the practical question is whether you can export a defined subset of the archive with its metadata and integrity manifests intact. If the answer requires reconstructing which files belong to which block, that work will be done under transaction deadline pressure.
Partner and joint-venture data obligations
Most upstream assets are operated jointly, and the data has more than one interested party. That creates requirements a single-owner archive never encounters.
- Partition by asset, not by department. Non-operated interests and operated blocks need to be separable without a manual sift.
- Provable non-tampering. Where partners are entitled to data supporting a cost or reserves position, checksums and write evidence answer questions that assurances do not.
- Export packages, not access grants. Partner delivery is usually a self-describing package with a manifest, not a login to your archive.
- Confidentiality windows. Some data becomes shareable or releasable only after a defined period, which is a retention attribute in its own right.
Legacy tape from earlier decades
Every long-lived upstream estate has a shelf of media from formats whose drives are no longer manufactured. Nine-track reels, early cartridge formats, and generations of LTO that current drives will not read. Read compatibility is limited to a small number of generations back, so an archive left untouched for long enough becomes unreadable through simple obsolescence rather than through any media failure.
Recovery is possible more often than teams expect, but it is specialist work with a real failure rate, and the honest planning assumption is that some percentage of very old media will not be recoverable. The way to avoid the problem is a scheduled refresh cycle with verification, described under tape and LTO systems. Aban Smart designs, supplies, integrates and supports archive platforms from QStar, INCOM StorEasy and DISC; we do not operate a media recovery laboratory.
Adjacent estates with similar object sizes include research and HPC and media and video. See the industries overview, or request an assessment with an inventory of your oldest media formats.
Frequently asked questions
That is a commercial judgement, but discarding raw records is irreversible and removes the option to reprocess with future algorithms. The usual compromise is to hold raw data on the cheapest durable tier with a slow retrieval service level, and keep final volumes on faster storage. Deleting the intermediate products is normally the safer saving, since they can be regenerated from raw data.
Related pages
Turn your requirement into a defensible architecture
Share the workload, capacity, retention, access, and resilience requirements. Aban Smart will identify the next discovery inputs and the appropriate engagement path.