If your customer files live on a file server, a NAS or a shared drive, the script on this page copies them into Files with their folder structure. Run it again whenever you like: files that haven't changed are recognised without being sent, so later runs only upload what is new or changed. It uses the Files API and nothing else, so you can read it, change it, or use it as the starting point for your own integration.
Before you start#
- Python 3.10 or later and the
requestspackage, on a computer that can read the archive. A wired connection helps: a large archive can take many hours, so run it overnight. - An API key (see the API overview). The key acts as its user, and every version it uploads shows that key's name in the file's activity. A dedicated user for the import keeps that clear.
- Enough storage. Run the script with
--dry-runfirst: it reports how many files are new, how many become new versions, how many are unchanged, and how much will be sent, without uploading anything. See Storage.
Run it#
Save the script below on that computer, then:
python3 -m pip install requests
export API_KEY="paste your key here"
python3 files_sync.py "/Volumes/Archive/Customers" --host https://app.example.com --team acme --into Customers --dry-run
python3 files_sync.py "/Volumes/Archive/Customers" --host https://app.example.com --team acme --into Customers
| Option | What it does |
|---|---|
| the folder | The folder to copy. Its contents go into the --into folder; the folder itself isn't created again. |
--host |
The address you open the app at. |
--team |
Your Team ID (see Teams, members and roles). |
--into |
A folder path in Files, such as Customers or Archive/2025. Missing folders are created. Leave it out to copy into the top of Files. |
--dry-run |
Only report what would be uploaded. |
--properties |
A spreadsheet of property values to set after the upload (see below). |
What the script does#
- Walks the folder, skipping the same system files as the web app (see Folders and uploads) and hidden folders.
- Creates every folder in Files, empty ones included. Folders that already exist are reused.
- Plans the upload in batches of 1,000 files. A file with the same size and modified date as the current version in Files is unchanged and isn't sent. A file whose name already exists becomes that file's next version (see Versions).
- Sends the bytes straight to storage, six files at a time. Files of more than 64 MiB go in parts.
- Completes the uploads in batches of 50, then waits until Files has checked every file, and prints the counts: new files, new versions, unchanged, skipped and failed.
- Sets properties from your spreadsheet, if you gave one.
Every version keeps the file's path in the archive and its modified date, which Files shows in the version's details.
Running it again#
- Only changes are sent. New files are added, changed files become new versions, and unchanged files are left alone.
- Nothing is deleted. A file you deleted or renamed in the archive stays in Files; a renamed file arrives as a new file.
- Interrupted runs resume. If the connection drops or the computer sleeps, run the same command again within 48 hours. The parts of large files that already arrived aren't sent again.
Properties from a spreadsheet#
Save a CSV file with a path column, the file's path inside the folder you copy, and one column per property
key (see Properties and tags):
path,customer,sku,due_date
C-102/Postcard A.pdf,C-102,PC-A,2026-10-15
C-102/Menu 11x17.pdf,C-102,MN-11,
- Empty cells are skipped, so they never clear a value.
- Values are checked against each property's type. Dates are
YYYY-MM-DD, yes / no values aretrueorfalse, and a choice must be one of its choices. Multiple-choice properties can't be set from the spreadsheet. - Rows whose file isn't in Files, and values that are refused, are printed, and the other rows still apply.
When every file in a folder shares a value, such as its customer, a folder default is simpler: set it once on the folder instead.
The script#
#!/usr/bin/env python3
"""Copy a folder tree into Files, keeping its structure. Run it again to send only what changed.
python3 -m pip install requests
export API_KEY="paste your key here"
python3 files_sync.py "/Volumes/Archive/Customers" --host https://app.example.com --team acme --into Customers
Options: --dry-run (report only), --properties values.csv (a "path" column and one column per property key).
"""
import argparse
import csv
import hashlib
import json
import os
import random
import sys
import time
import uuid
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from urllib.parse import urljoin
import requests
JUNK = {".DS_Store", ".localized", "Thumbs.db", "desktop.ini", "Icon\r"}
PLAN_BATCH, COMPLETE_BATCH, ENSURE_BATCH, THREADS = 1000, 50, 2000, 6
RETRY = {408, 429, 500, 502, 503, 504}
def walk(source):
"""Every folder below source (empty ones too) and every file, as paths relative to source."""
folders, files = [], []
for here, dirs, names in os.walk(source):
dirs[:] = sorted(d for d in dirs if not d.startswith("."))
rel = Path(here).relative_to(source).as_posix()
if rel != ".":
folders.append(rel)
for name in sorted(names):
if name not in JUNK and not name.startswith(("._", "~$")):
files.append(Path(here, name).relative_to(source).as_posix())
return folders, files
class Files:
def __init__(self, host, team, key):
self.api = "/".join([host.rstrip("/"), "a", team, "files", "api", "v1", ""])
self.headers = {"Authorization": f"Api-Key {key}"}
def call(self, method, path, allow=(), headers=None, **kwargs):
"""One API call, retried when the server is busy. Exits on an error, except the statuses in allow."""
for attempt in range(6):
response = requests.request(
method, self.api + path, headers={**self.headers, **(headers or {})}, timeout=300, **kwargs
)
if response.status_code not in RETRY:
break
time.sleep(min(60, 2**attempt))
if response.status_code in allow:
return None
if response.status_code >= 400:
sys.exit(f"{method} {path}: {response.status_code} {response.text[:500]}")
return response.json()
def sign(self, upload_id, parts=None):
item = {"id": upload_id, "parts": parts} if parts else {"id": upload_id}
return self.call("POST", "uploads/sign/", json={"items": [item]})["items"][0]
def transfer(http_method, url, body, fields=None):
"""Send bytes to storage. Returns False when the link has expired and must be signed again."""
for attempt in range(5):
try:
if http_method == "POST":
response = requests.post(url, data=fields, files={"file": body}, timeout=900)
else:
response = requests.put(url, data=body, timeout=900)
if response.ok:
return True
if response.status_code == 403:
return False
except requests.RequestException:
pass
time.sleep(2**attempt + random.random())
raise RuntimeError(f"storage refused the upload after 5 tries: {url[:60]}...")
def send(files, path, upload):
"""One file: a single POST or PUT, or the parts storage doesn't have yet (so a re-run resumes)."""
if not upload.get("part_count"):
body = path.read_bytes()
while not transfer(upload["http_method"], urljoin(files.api, upload["url"]), body, upload.get("fields")):
upload = files.sign(upload["id"])
return
count, size = upload["part_count"], upload["part_size"]
have = {p["number"] for p in files.call("GET", f"uploads/{upload['id']}/parts/")["parts"]}
urls = {p["number"]: urljoin(files.api, p["url"]) for p in upload["parts"]}
with path.open("rb") as fh:
for number in (n for n in range(1, count + 1) if n not in have):
fh.seek((number - 1) * size)
body = fh.read(size)
while True:
if number not in urls:
wanted = [n for n in range(number, min(count, number + 19) + 1) if n not in have]
urls.update(
{p["number"]: urljoin(files.api, p["url"]) for p in files.sign(upload["id"], wanted)["parts"]}
)
if transfer("PUT", urls[number], body):
break
del urls[number]
def plan_items(source, prefix, chunk):
items = []
for rel in chunk:
stat = (source / rel).stat()
handle = hashlib.sha256(f"{rel}|{stat.st_size}|{stat.st_mtime_ns}".encode()).hexdigest()[:40]
items.append(
{"client_id": handle, "relative_path": prefix + rel, "client_path": rel, "size": stat.st_size,
"last_modified": int(stat.st_mtime * 1000)}
) # fmt: skip
return items
def set_properties(files, prefix, csv_path):
"""Rows of a CSV (path, then one column per property key): one bulk call per distinct set of values."""
groups = {}
with csv_path.open(newline="", encoding="utf-8-sig") as fh:
for row in csv.DictReader(fh):
path = "/" + prefix + row.pop("path").strip("/")
found = files.call("GET", "resolve/", params={"path": path}, allow=(404,))
if found is None or found.get("type") != "file":
print(f" not in Files: {path}")
continue
values = {key: value for key, value in row.items() if key and value and value.strip()}
groups.setdefault(json.dumps(values, sort_keys=True), []).append(found["id"])
for values, ids in groups.items():
for i in range(0, len(ids), 1000):
body = {"action": "set_properties", "file_ids": ids[i : i + 1000], "properties": json.loads(values)}
result = files.call("POST", "bulk/", json=body, headers={"Idempotency-Key": str(uuid.uuid4())})
for item in result.get("results") or []:
if item["status"] == "error":
print(f" file {item['id']}: {item['error']['detail']}")
def main():
parser = argparse.ArgumentParser(description="Copy a folder tree into Files.")
parser.add_argument("source", type=Path, help="the folder to copy")
parser.add_argument("--host", required=True, help="the address of the app, e.g. https://app.example.com")
parser.add_argument("--team", required=True, help="your Team ID, e.g. acme")
parser.add_argument("--into", default="", help="a folder path in Files, e.g. Customers (default: the root)")
parser.add_argument("--properties", type=Path, help="a CSV of property values to set after the upload")
parser.add_argument("--dry-run", action="store_true", help="only report what would be uploaded")
args = parser.parse_args()
files = Files(args.host, args.team, os.environ["API_KEY"])
into = args.into.strip("/")
prefix = f"{into}/" if into else ""
folders, paths = walk(args.source)
print(f"{len(paths)} files in {len(folders)} folders")
if not args.dry_run:
wanted = ([into] if into else []) + [prefix + folder for folder in folders]
for i in range(0, len(wanted), ENSURE_BATCH):
files.call("POST", "folders/ensure/", json={"paths": wanted[i : i + ENSURE_BATCH]})
batch, totals, needed, uploads = str(uuid.uuid4()), {}, 0, []
for i in range(0, len(paths), PLAN_BATCH):
chunk = paths[i : i + PLAN_BATCH]
body = {"batch": batch, "dry_run": args.dry_run, "items": plan_items(args.source, prefix, chunk)}
plan = files.call("POST", "uploads/", json=body)
needed += plan["quota"]["needed"]
for rel, item in zip(chunk, plan["items"], strict=True):
totals[item["status"]] = totals.get(item["status"], 0) + 1
if item["status"] == "error":
print(f" {rel}: {item['error']['detail']}")
elif "upload" in item:
uploads.append((args.source / rel, item["upload"]))
print("plan:", totals, f"{needed / 1e9:.2f} GB to send")
if args.dry_run:
return
with ThreadPoolExecutor(THREADS) as pool:
for i in range(0, len(uploads), COMPLETE_BATCH):
chunk = uploads[i : i + COMPLETE_BATCH]
list(pool.map(lambda job: send(files, *job), chunk))
items = [{"id": upload["id"]} for _path, upload in chunk]
done = files.call("POST", "uploads/complete/", json={"items": items},
headers={"Idempotency-Key": str(uuid.uuid4())}) # fmt: skip
for item in done["items"]:
if item["status"] == "error":
print(f" {item['name']}: {item['error']['detail']}")
print(f" sent {i + len(chunk)} of {len(uploads)}")
while uploads:
status = files.call("GET", "uploads/", params={"batch": batch})
if status["settled"]:
print("done:", status["counts"])
break
time.sleep(10)
if args.properties:
set_properties(files, prefix, args.properties)
if __name__ == "__main__":
main()
When something goes wrong#
| The script prints | What to do |
|---|---|
403 |
The key is wrong or revoked, or its user isn't a member of the team in --team. |
507 and quota_exceeded |
The files don't fit in your storage. Nothing was reserved. See Storage. |
name_invalid for a file |
The name can't be used in Files (see Folders and uploads). Rename it in the archive. |
too_large for a file |
The file is over the 10 GiB limit. |
upload_incomplete |
The bytes didn't all arrive. Run the script again. |
Busy servers (429 and 503) are retried automatically, waiting a little longer each time.