QStar and BeeGFS Solution Brief
Archiving BeeGFS parallel file-system data for HPC and research workloads.
- Vendor:
- QStar / BeeGFS
- Date:
- Jun 2026
- Format:
- Size:
- 1.1 MB
Written now, read in 2040
The next person to open it is a reviewer, a doctoral student or a retraining job, working from a citation and whatever survived of the metadata.
Aban Smart designs, supplies, integrates and supports archive software sitting behind parallel filesystems and research computing platforms across the GCC.
Research storage has an audience problem. Whoever opens a dataset next is usually a stranger to it: a reviewer in year three, a student in year eight, a retraining run in year twelve. None of them can ask the postdoc what final_v2_use_this held, because she left and the cluster she used was decommissioned twice over. Every requirement below follows from that. Ingest speed matters least. Whether an outsider can locate, trust and reuse the data long after the grant closed matters most.
Datasets are re-read years later by people who did not create them.
Facilities converge on three tiers, and most complaints about research storage are really data sitting in the wrong one.
| Tier | Sits on | Contains | Expected life |
|---|---|---|---|
| Scratch | Parallel filesystem, NVMe or fast disk | Job working sets, checkpoints, intermediate output | Days to weeks, then purged |
| Project | Group-quota'd filesystem | Curated inputs, code, working results | Duration of the grant |
| Archive | Tape, optical, object or cloud | Reduced outputs, deposited datasets, training corpora | Ten years upward |
Scratch is not a small archive. It is a fast, expensive job workspace with a purge schedule attached, and letting it quietly become permanent ruins a facility's economics.
An illustration, with invented figures. Substitute your own instrument or job logs before deciding anything.
Take a beamline or simulation campaign producing 4 TB per run, executed 250 times a year.
That ratio is the budget conversation. Archive all of it and the archive is five times larger for no research gain. Archive none of it and 1 PB of irreplaceable output waits on the costliest storage on site until something deletes it. The percentage varies sharply by discipline, so derive it per group.
Scratch fills, the queue backs up, a purge is announced, deferred, softened, then enforced badly. It repeats because the two sides carry different risks: the facility pays for the capacity, the research group loses the work. Policy statements alone have never fixed this. Making the exit path easier than doing nothing does.
Ownership of the purge policy stays with the institution. The storage layer's contribution is making the compliant choice the low-effort one.
A parallel filesystem is tuned for concurrent throughput, not cold data, and its metadata services degrade first when it fills with files nobody has touched in two years. The design question is how data exits without a gateway becoming the choke point.
QStar announced an integration with the BeeGFS parallel filesystem in June 2026. Per that announcement, QStar Network Migrator, a hierarchical storage management product, uses the BeeGFS Data Management API to read file metadata directly instead of walking the tree, which the vendor states reduces scan time and impact on primary storage. Policies can be written against last access time, ownership, group membership, size or file type. Files move while access continues through lightweight links or stubs.
Underneath, QStar Archive Manager supplies NFS archive gateways with caching, targeting tape libraries, object storage and cloud. For shared national or multi-faculty facilities, QStar's Global ArchiveSpace datasheet describes a clustered active archive with a global namespace, front-end access over SMB, NFS, S3, HTTP and FTP, and a Physical Library Manager that pools tape drives across nodes rather than partitioning the library. One inconsistency in that document: the highlights state 3 to 64 nodes while a diagram shows 2 to 64. Confirm the floor before sizing. See tape and LTO systems for the media layer underneath.
A citation is a promise that an identifier still resolves years after the storage under it was replaced. Routine infrastructure work breaks that promise quietly. Three things hold it together.
Funders and institutions commonly attach availability conditions to grants, and those conditions outlast the hardware. Nothing here states what any funder, regulator or policy requires of you. That reading belongs to your research office and legal advisers. Our concern is that a resolution request in year twelve returns the right object.
Model reproducibility is stricter than dataset citation because the answer must be exact. "Trained on the genomics corpus" is not an answer. "Trained on this frozen snapshot, manifest hash included, on this date" is.
Write-once snapshots are designed to prevent a corpus drifting under a model. They say nothing about whether the model is accurate, fair or safe, and no storage design can.
Research capability across the region is being built quickly, and compute procurement moves faster than data policy. A cluster is specified, funded and installed while the archive is still a future phase. The gap surfaces around eighteen months in, when scratch is full and the first finished projects have nowhere to go.
Two consequences follow. The archive is then sized against data that already exists rather than a growth model, so it arrives undersized. And files deposited during the gap carry no consistent metadata, which is expensive to retrofit.
Adjacent problems appear in oil and gas, where decades-old survey data must stay re-ingestible, and in media and surveillance, where the archive is sized by ingest rate. Browse the other industries we work in, or request an assessment with your annual production figures and current scratch policy.
Manufacturer documentation relevant to this page. Availability, specifications, and configurations are subject to verification.
Archiving BeeGFS parallel file-system data for HPC and research workloads.
Global namespace across distributed archive locations for HPC, AI, and media workloads.
Announcement of QStar and BeeGFS integration for research data archiving.
That is the point of a DMAPI-based HSM layer. QStar's BeeGFS integration reads metadata through the filesystem's data management interface rather than scanning directories, which the vendor material presents as cutting scan time and load on primary storage. Data movers then write to archive destinations in parallel. Validate the claim against your own file-count profile and job mix before you commit to a design.
Share the workload, capacity, retention, access, and resilience requirements. Aban Smart will identify the next discovery inputs and the appropriate engagement path.