How to migrate to a DAM without losing ten years of files
A DAM migration is a curation exercise, not a file transfer. Here's the phased plan: what to move, what to leave behind, and how to make people actually switch.
There's a moment in every DAM migration when someone opens the shared drive to see what's actually in there, and goes quiet. Ten years of folders. Three naming conventions. A directory called FINAL_final_v2_USE THIS ONE. This is where most projects go wrong, in one of two directions.
Either the team tries to move everything, the audit balloons under the weight of 300,000 files, and the platform sits half-populated for a year. Or the team moves everything unexamined, and the new system is the old mess with a better search bar.
The way through is accepting one thing up front: a DAM migration is a curation exercise, not a file transfer, and most of your library doesn't deserve the effort.
Phase 0: decide what "done" means
Before you touch a single file, write down the finish line. Something concrete:
By 1 November, all campaign photography from 2023 onward, the complete brand kit, and all approved video is in the DAM with rights status and campaign recorded. The old drive is read-only.
Two things make this a good goal. It's bounded: a date and a scope, not "our assets". And it has a switch-off: the old drive goes read-only, which is the only mechanism that reliably changes behaviour.
Assign one owner. Not a committee. A person whose name is on the outcome.
Phase 1: audit, honestly
You need four numbers and one map.
The numbers: total assets, total size, the share last opened more than two years ago, and the share that are near-duplicates. Those last two shock most teams. A library where 70% of the files haven't been touched in two years is completely normal, and knowing it reframes the whole project.
The map: every place assets actually live. Not the official list. The real one. The shared drive, the SharePoint sites, three Dropbox accounts, the marketing manager's external hard disk, the agency's WeTransfer archive, the photographer's own Lightroom catalogue, the old DAM nobody logs into anymore. Write them all down, each with an owner.
The shadow locations are the ones that matter. A migration that leaves the photographer's hard drive as the real archive hasn't migrated anything.
Phase 2: triage into four buckets
Go through the map and sort. Rough is fine; you're trying to shrink the problem, not perfect it.
| Bucket | What goes in it | What happens to it |
|---|---|---|
| Move and enrich | Assets in active use, the brand kit, the last two to three years of campaigns, evergreen photography | Full treatment: metadata, rights, approval |
| Move as-is | Older material with residual value, historical archive | Bulk import into an archive area; searchable, minimal metadata |
| Leave | Working files, drafts, superseded exports, personal copies, anything with no reuse value | Stays in cold storage or gets deleted |
| Decide later | Genuine unknowns | A single "to triage" area with a review date, and a hard rule that it is not a permanent home |
Expect "move and enrich" to be 10–20% of your library. That's the number that makes this project finishable.
The instinct to keep everything is strong, and it's mostly wrong. Assets with no metadata, no known rights and no recent use aren't an archive. They're a liability with storage costs. If deleting feels impossible, move them to cheap cold storage outside the DAM and write down where they went.
Phase 3: design the metadata model first
Do this before the first bulk upload. Retro-fitting a schema across 40,000 assets is painful, and doing it twice is worse.
Keep it small: four or five required fields, a handful of optional ones, controlled vocabularies on anything you'll filter by. The full method lives in metadata and taxonomy, but three pieces of advice are migration-specific.
Mine your folder names. They're a metadata model already, written by the very people who used the library. 2024-09_OpenDay_Hasselt/Photography/Selects/ yields a date, an event, a location, an asset type and an approval state. Map the folder patterns to fields once, apply them in bulk, and most of your descriptive metadata comes for free.
Mine EXIF and IPTC. Capture date, camera, GPS, plus whatever credit or copyright the photographer embedded. Free, accurate, and already sitting inside the file.
Don't recreate the folder tree. This is the most common migration mistake, and it always starts innocently. The tree comes across as collections or workspaces "just so people can find their way", it becomes the primary navigation, and six months later you've rebuilt the exact system you were escaping. If you genuinely need a transitional structure, give it an end date.
Phase 4: migrate in tranches
Never in one pass. Four or five tranches, each with its own checkpoint.
- The brand kit. Small, high value, everyone needs it, and it proves the system works. Logos, colours, fonts, templates, guidelines.
- The current campaign. Live work, which forces daily use from day one.
- The last 12–24 months. The bulk of your "move and enrich" bucket.
- The historical archive. Bulk import, minimal metadata, straight into the archive area.
- The shadow locations. The hard drives and personal accounts, chased down one owner at a time.
After each tranche, stop and look. Are people actually using it? Is search returning what it should? Is the metadata getting filled in? A problem you catch after tranche one is a conversation. The same problem after tranche five is a re-migration.
Deduplicate on the way in
Let the platform catch exact and near-duplicates at ingest; it's far cheaper than cleaning up afterwards. Decide the rule in advance (usually: keep the highest resolution and the earliest capture date) and record the others as variants rather than deleting them silently.
Preserve what you can't regenerate
Original files, always. Embedded IPTC and XMP. Capture dates (sort a decade of photography by upload date instead, and it all looks like it happened last Tuesday). And keep the mapping from old path to new asset, so when someone digs up a link from an old email, you have an answer.
Phase 5: enrich where it pays
You won't enrich everything, and you shouldn't try.
Let AI take the first pass. Descriptions, objects, scenes, text-in-image and face clustering, run across the whole library in bulk. It makes even the unenriched archive semantically searchable. One caveat: check what it costs at your volume first, because some platforms meter AI processing per asset. AI search in a DAM covers what these features actually do, and where they stop.
Humans add only what AI can't know. Campaign, rights status, approval, project codes. Concentrate that effort on the "move and enrich" bucket.
Enrich on retrieval. For everything else, prompt for the missing field when someone actually opens the asset. The effort lands exactly where the value is, and the person retrieving it is the one who knows the context.
Set a rights baseline. Anything whose licence you can't establish gets marked rights unknown, do not publish, never left blank. Blank reads as permission. It's a fifteen-minute decision that prevents a genuine incident.
Phase 6: switch over
The technical work is done. This is the part that decides whether the project actually worked.
Make the old location read-only on a named date. Announce it three weeks out, remind at one week, then do it. A migration with no switch-off produces two libraries indefinitely, and people will keep using the one they know.
Redirect the requests, not just the files. The colleague who's always asked Sarah for photos will keep asking Sarah. Sarah's job for the first month is to answer with a link into the DAM every single time, never with a file. Two or three rounds of that and the habit moves.
Train in ten minutes, not two hours. For 90% of users, the entire training is: here's the URL, type what you're looking for, click download. Save the long session for the people who upload and configure.
Watch the search logs. Queries that return nothing are the most valuable feedback you'll get. They tell you exactly which vocabulary is missing, which assets never got enriched, and what people expected to find that you don't have. Review them weekly for the first two months.
Name the fallback. Someone will need a file that didn't make the cut. Publish where the cold archive lives and who can retrieve from it, so the answer is never "I think it's on the old drive somewhere".
Migrating from a legacy DAM
A different problem: slightly easier, and slightly worse.
Easier: you already have metadata, and it usually exports.
Worse: that metadata is shaped for the old system, and its taxonomy encodes years of accumulated compromise. Don't import it verbatim. Map the old fields to your new model deliberately, and treat the migration as your one chance to merge the duplicate terms nobody was allowed to touch.
Two practical warnings. First, confirm you can export original files and metadata together, in a documented format, before you sign anything with the new vendor. Second, check whether your existing shared links will break: if the old system has public URLs embedded in a website, a press release or printed material, you need a redirect plan.
A realistic timeline
For a marcom team with 50,000–150,000 assets:
| Week | Focus |
|---|---|
| 1 | Audit, map every location, agree the finish line |
| 2 | Triage into four buckets; design the metadata model |
| 3 | Tranche 1 (brand kit) and tranche 2 (current campaign); first users onboarded |
| 4–5 | Tranche 3 (last 24 months) with bulk enrichment |
| 6 | Tranche 4 (archive), bulk AI pass |
| 7 | Shadow locations; rights baseline pass |
| 8 | Old drive read-only; review search logs; hand over ownership |
Eight weeks, of which roughly two are actual file movement. The rest is decisions.
Frequently asked questions
How long does a DAM migration take?
Four to eight weeks for a typical marketing team, and most of that time goes to triage, metadata design and enrichment rather than moving files. Enterprise migrations with multiple source systems and integrations run longer.
Should we migrate all our assets?
No. Audit first: most libraries carry a large tail that hasn't been opened in years. Fully enrich the 10–20% in active use, bulk-import the rest into an archive area with minimal metadata, and leave working files and superseded exports behind.
Can we keep our folder structure in the DAM?
You can, and it usually undermines the project. Folder trees force a single classification onto every asset, and that's the very limitation you're migrating away from. Use folder names as a source of metadata during import, then let search and filters replace the tree.
How do we get people to actually use the new system?
Make the old location read-only on an announced date, answer every asset request with a link into the DAM rather than a file, and keep training to ten minutes for people who only download. Adoption follows search quality more than it follows training.
What should we do with assets whose rights we cannot verify?
Mark them explicitly as rights unknown and block publication rather than leaving the field blank, because blank gets read as permission. Verify the ones that matter, and archive the rest out of the default view.