Blog

Excel metadata: device fingerprints in every spreadsheet

Excel metadata: device fingerprints in every spreadsheet

Two analysts in different companies open the same spreadsheet. They see the same numbers, the same charts, the same formatting. But inside the file, hidden from the cells themselves, sits a small block of Excel metadata that uniquely identifies who first authored it, what computer they used, which template it descended from, and a per-session revision ID for every block of cells that has ever been edited. Two copies of the same workbook, used in different organisations, will not have the same fingerprint.

Most of this is the same plumbing as Word. A few pieces are specific to Excel and worth understanding before you forward a workbook outside the firm.

The Office Open XML structure

An .xlsx is a ZIP. Inside, the same docProps/core.xml and docProps/app.xml hold the high-level properties: creator, lastModifiedBy, company, application version, total editing time, last printed timestamp. None of that is exposed inside the cell grid; you have to unzip the file or run a metadata tool to see it.

Underneath, in xl/worksheets/sheet1.xml, every cell carries its formula or value. Unlike Word, Excel does not track per-paragraph revisions, but it does store something arguably more revealing.

Revision save IDs and per-session fingerprints

Each save event writes a fileVersion entry and, in older builds, a per-revision rsid table similar to Word's. More usefully for a forensic reader, Excel attaches a headerFooterChanges log when those have been edited, and records the cell range affected by each change. Combined with the lastModifiedBy field, that lets a recipient reconstruct, in broad strokes, who edited which sheet of a financial model and roughly when.

If the workbook has Track Changes enabled (or has had it enabled in the past, in a shared workbook), an entire revision log lives in xl/revisions/revisionLog.xml. It records cell-by-cell changes with the user account that made them, in order. Forwarding a tracked workbook outside the company without removing that log can hand the recipient a complete edit history.

The hidden things people forget

Defined names and external links

Workbooks built from a template often hold defined names that point to ranges in source files: '\\fileserver01\Finance\Models\BaseAssumptions.xlsx'!Inputs. Even if the data has been pasted as values, the named range can persist. The recipient sees the network path of the originating file, sometimes including a server name and folder structure that the company would never publish.

Comments and notes with author names

Cell comments (called Notes in newer builds) are stored alongside the author's display name as configured in their copy of Excel. A note saying "Check this with Sarah" attached by Tom appears in the file as author="Tom Wilson". Threaded comments — the newer style — carry full Microsoft 365 user identifiers (UPN), which is a real email address.

Pivot caches

A pivot table built on data from a much larger source caches a copy of the original dataset inside the workbook. The visible pivot might show 12 anonymous rows; the cache underneath can hold tens of thousands of unredacted entries. Sending the workbook to an outside reader without clearing the cache hands them the full source data.

Hidden and very-hidden sheets

Excel has two levels of sheet hiding. "Hidden" can be unhidden through the right-click menu. "Very hidden" sheets do not appear in that menu and can only be shown via VBA or by editing the file's XML directly. Drafts, working notes and assumption tables tucked into very-hidden sheets often travel with the workbook unintentionally.

Embedded objects

The same problem as Word: pasted Outlook screenshots, embedded PowerPoint charts and OLE objects can carry their source files inside the workbook.

Macros and VBA project metadata

An .xlsm with a Visual Basic project carries the project's name, the module list, the digital signature of the signer, and frequently the original developer's username embedded in the binary stream. Even a workbook whose macros do nothing visible can carry a developer attribution that the recipient never asked for. The same applies to Power Query and Power Pivot models: connection strings can include a server name, a database name and sometimes the credentials path used to authenticate against the source system.

What to clean before you send a model out

Excel's Document Inspector handles document properties and the obvious comment author names. It does not always touch defined-name external references, pivot caches with un-anonymised source rows, or very-hidden sheets you did not know were there. For workbooks going outside the organisation — investor decks, client deliverables, regulatory filings — the right move is a deliberate three-step pass: clear pivot caches and externalised links inside Excel, run the Inspector, then strip the residual document and revision metadata. The cleaner on the homepage handles the last step, including the .xlsm and .xltx variants the Inspector skips.

The point is not that Excel is leaky by design. The metadata is genuinely useful while you are working on the file. It just stops being useful, and starts being a small dossier on your team, the moment the file leaves your network.