Working with Zip Files
James Carr
September 28th, 2026
ZIP files are handy. They bundle many files into one neat package that's easy to email, download or move between systems. But a ZIP is built for transfer, not for long-term preservation. When it comes to digital preservation, storing ZIP files in a preservation system such as Preservica is generally not seen as good practice.
A digital preservation system works best when it can see what it's looking after. Wrapped inside a ZIP, your content is effectively hidden. The system recognises a single archive file, not the documents, images, spreadsheets or emails inside it.
The problem with keeping files zipped
- The contents are hidden. Format identification (e.g. DROID), fixity checks at file level, validation and metadata extraction only see “a ZIP”. You can’t tell whether it holds PDFs, obsolete WordPerfect files or executables.
- Preservation actions can’t reach the files. You can’t migrate a format or plan for obsolescence on individual files you can’t see.
- One damaged byte can cost you many files. Corruption in the central directory or a compressed stream can make several files, or the whole archive, unreadable. Uncompressed files are far more resilient to errors.
- It’s harder to find and use. You can’t describe, search, render or give access to individual items.
Sometimes you will need to store a Zip file, for example the file is technically a Zip underneath but has its own documented format, such as DOCX, XLSX, EPUB and ODF.
The Zip format is used frequently as a package transport format, for example zipped Bagit packages etc. The Zip format here should only be used for transport and delivery, never as the stored archival format.
Having said all that, sometimes Zip files cannot be avoided, when the Zip is the record itself, for example a software release, a dataset exactly as published, or evidence where the original package matters.
In the case where a Zip file has been stored, accessing the individual files inside the Zip stored in Preservica requires downloading the entire Zip file and extracting its contents to get to the file of interest. For large Zips this could mean downloading many GBs of data to access a small file inside the Zip.
The 3rd party Python SDK pyPreservica has now added two new functions to allow users to interact with Zip files without having to download them out of the Preservica repository.
The first is a function to allow users to list the contents of a Zip file held in Preservica.
from pyPreservica import *
client = EntityAPI()
asset = client.asset("9fd239eb-19a3-4a46-9495-40fd9a5d8f93")
for bs in client.bitstreams_for_asset(asset):
for name in client.bitstream_zip_names(bs):
print(name)
The function takes a BitStream object and returns a list of file names within the Zip container. You can use this to enumerate all the compressed objects within the Zip file.
This Python script will print the name of every Zip entry to the console without having to download the Zip file. You can use this to quickly check the contents of the Zip. This call is very quick as only parts of the Zip containing the header information is accessed.
If your Asset does not contain a Zip file, then nothing will be printed.
Once you have the name of the file inside your stored Zip, you can export it out of Preservica directly without having to download the Zip and extract it locally on your machine.
For example, if you have a Preservica Asset which contains a very large Zip file containing multiple video files and a small XML Mets file, if you only need access to the XML, you can extract it locally using:
from pyPreservica import *
client = EntityAPI()
asset = client.asset("9fd239eb-19a3-4a46-9495-40fd9a5d8f93")
for bs in client.bitstreams_for_asset(asset):
client.bitstream_zip_content("mets.xml")
This will write the contents of the mets.xml file into a local file in the current directory. Again, because the Zip file is not downloaded and only the bytes in the requested file are read from the server, this will be a very fast call.
More updates from Preservica
Custom Reporting via the Preservica Content API
Preservica provides a REST API to allow users to query the underlying search engine. In this article we will show how CSV documents can be returned by the API.
James Carr
November 29th, 2021
Using OPEX and PAX for Ingesting Content
Preservica has developed the concept of an OPEX (Open Preservation Exchange) package, a collection of files and folders with optional metadata, as a way to organise content into an easy to understand format for transfer into or out of a digital preservation system. Although we have created it, we hope suppliers of digital content to be preserved, and other digital preservation systems, will use it due to its simplicity.
Richard Smith
January 28th, 2021
Using the PAR API to create Custom Migrations
Since the release of v6, Preservation Actions within Preservica have been defined and controlled using a PAR (Preservation Action Registries) data model. To facilitate this, Preservica’s registry also exposes a PAR API to allow a full range of CRUD operations on this data. This API also makes it possible to write new migration actions using Preservica’s existing toolset, for example, to introduce re-scaling to your image/video migrations, or to get different output formats altogether. In this article, we will introduce the key concepts in this data model, explain how Preservica uses and interpret them, and introduce the API calls required to create your own custom actions. We will do this by a worked example, using ImageMagick to create a custom “re-size migration” for images.
Jack O'Sullivan
August 11th, 2020
Using Python with the Preservica Entity APIs (Part 3)
In this article we will be looking at API calls which create and update entities within the repository, some calls to add and update descriptive metadata and we will also look at the use of external identifiers which are useful if you want to synchronise external metadata sources to Preservica.
James Carr
June 11th, 2020
Preservica on Github
Open API library and latest developments on GitHub
Visit the Preservica GitHub page for our extensive API library, sample code, our latest open developments and more.
Preservica.com
Protecting the world’s digital memory
The world's cultural, economic, social and political memory is at risk. Preservica's mission is to protect it.