How a Flow Works — the OpenFlow YAML, X-Rayed
A flow looks like a diagram. Underneath it is just a YAML file — learn to read one and you read them all.
Plain-English, diagram-rich explanations of how Windmill actually works.
Got your own code? Open the Explainer →37 concepts · the complete catalog · press / to search
A flow looks like a diagram. Underneath it is just a YAML file — learn to read one and you read them all.
You can do everything in one script — so when do you reach for a flow? The answer is about what the platform can see.
Your script needs a database password — you will not type it into the script. Here is the Windmill way.
A Windmill Variable and a Secret look like the same thing — and the difference is one checkbox that matters a lot.
When you hit Run, your code does not run where you are sitting — and that one fact explains the most confusing bugs in Windmill.
In a flow, step five can use the output of step one — not because the data trickled down, but because step five reached back and took it.
Two Windmill constructs are both called 'branch' — one takes a single path, the other takes them all, and picking wrong is a classic bug.
You have 200 files to process — you do not write the processing 200 times, and you describe the loop just once.
Two constructs decide 'should the flow keep going?' — one loops the body again, one stops everything; reason about them backwards and they bite.
Steps fail — a network blips, an API rate-limits — and Windmill gives you two different answers to 'what now?'.
Some steps do the same expensive work over and over; cache_ttl lets a step remember its own answer for a while.
If step B and step C don't need each other, why would they wait in line? In a flow, they don't have to.
Your instinct is one big script. Windmill rewards you for cutting it into small ones — the skill is knowing where to cut.
Your script does import pandas — on your laptop that just works; on the worker it works only if you said so.
Your script has HOST = 'prod-db.internal' near the top — it works, and it is also why you have three copies of the script.
Your script runs on a worker you cannot see — when something breaks at 2 a.m., the only witness is the log.
A cron line on a server runs your script on time — until the server reboots, or it silently stops and nobody notices.
You wrote a great helper function — you do not paste it into every script that needs it.
You never build the form. You declare what you need, and Windmill draws it for you.
Windmill draws the form for you — the default form is correct, and often a little raw; shaping it is what makes it usable.
Some decisions a machine should never make alone — an approval step puts a human in the loop, on purpose.
A flow gives the user a form and a Run button — sometimes that is exactly right, and sometimes the user needs a screen.
Your script already returns the answer. The last mile is letting somebody read it without logging into Windmill and clicking Run.
Your parameter has a default. The form says the field is required. Both of those are true at the same time, and the reason will change how you write every signature.
You declared three dependencies — your worker installed forty; the other thirty-seven are the ones that surprise you.
Your code is correct — you proved it runs on your laptop, and it fails on the worker; it might not be the code.
You built it, you saved it, you can see it in the list — and the URL says it does not exist.
A variable will hold your API key, your file path and your feature flag. Hand it a document and you find the wall.
It ran perfectly for eight months. Then the person who wrote it changed roles, and nothing worked — with no error anyone could find.
You pushed three files. The workspace had two hundred. Only three survived.
The one line whose job was "this runs anywhere" is the only line that runs in exactly one place.
The file is listed, with a usage count, on a page called Assets. The write that was supposed to create it failed twenty minutes ago.
The deploy went green. Every check passed. The job died three minutes later asking for a package nobody packed.
Your scheduled job failed at six in the morning. Was it the driver, the password, one table out of seven, or the folder? The job knew. It was never built to say.
Your dry run passed. It never touched the one part that breaks: the write.
They want one file, same name, replaced every week. The obvious way to do that is also the way to leave them nothing at all on a Monday morning.
The resource had the right-sounding name. The job ran green. The file had a quarter fewer rows than it should, and nothing anywhere turned red.
A flow looks like a diagram. Underneath it is just a YAML file — learn to read one and you read them all.
A Windmill flow looks like a diagram on a screen — boxes, arrows, a play
button. But that diagram is not the flow. The flow is a plain text file:
<name>.flow.yaml. The diagram is just a drawing of the file. Learn to read
the file and you can read every flow ever written — including the one that
broke at 2 a.m.
Think of a printed orchestral score.
The score — the YAML — is the source of truth. Every note, in order, on paper. The conductor's gestures — the visual editor — are a convenient performance of that score: expressive, but they vanish when the music stops. The orchestra — the Windmill worker — is what actually turns marks on paper into sound.
You can conduct without reading music: wave along, trust the players. But the moment one instrument plays a wrong note, the person who can read the score is the one who finds the exact bar where it went wrong. The visual editor lets you conduct. This article teaches you to read the score.
A .flow.yaml file has three blocks that matter:
summary / description — the human label. What shows up in the flow
list and at the top of the run page.schema — the questions the flow asks before it runs. This is the input
form; it gets its own article, Inputs Become Forms.value.modules — the heart: an ordered list of steps.Each entry in modules is one node on the canvas, and it has three parts:
id — a short handle (a, b, c, ...). Other steps refer to this
step's output by this id.value — what the step does. Its type is the verb: script (run
code), branchone / branchall (split), forloop (repeat). The full set
lives in the OpenFlow construct catalog.input_transforms — how the step gets its arguments. This is the
wiring. Each argument is either a static value or a small JavaScript
expression like results.a.count (take the count field from step a's
output) or flow_input.env (take it from the form).The order is not a setting you flip — the order is the list. And data does not "flow down the arrows" by magic: each step pulls exactly the inputs it declares, by name.
That last part is the secret most people miss. The arrows you see on the
canvas are input_transforms made visible. A step receives step a's output
because it asked for it — results.a.
Here is a complete two-step flow. Read it top to bottom like the score it is:
summary: Process inbound CSV files # the human label
schema: # the form Windmill draws
properties:
env:
type: string
enum: [test, prod] # -> renders as a dropdown
value:
modules: # the ORDERED list of steps
- id: a # step 1
value:
type: script
path: f/nsc/s10_list_inbound_files
input_transforms:
env:
type: javascript
expr: flow_input.env # pull 'env' from the form
- id: b # step 2 — runs after 'a'
value:
type: script
path: f/nsc/s20_process_csv
input_transforms:
files:
type: javascript
expr: results.a # pull step a's whole output
Step b runs after step a for one reason only: it comes after it in the
modules list.
The canvas hides things the YAML shows.
Three of those hidden things bite people constantly:
input_transforms are invisible on the canvas. You see an arrow from
a to b. You do not see whether b takes all of a, or just
a.count, or a.rows[0]. Only the YAML tells you. When a step receives
the wrong data, the YAML is where the answer is.modules list. Reordering is not cosmetic.retry, stop_after_if, and cache_ttl live on the module too — and
they barely show on the canvas. A flow can silently retry a step three
times, or stop early, and the only evidence is two lines of YAML.The canvas is the photograph; the YAML is the X-ray — and when a flow breaks, you need the X-ray.
You can do everything in one script — so when do you reach for a flow? The answer is about what the platform can see.
You can put your whole job in one script. Read the files, clean the data, load
the database, send the email — one main(), top to bottom. Windmill lets you
do exactly that, and it will run fine. So why would you ever split it into a
flow?
The honest answer is not "big jobs need flows." Plenty of big jobs are one script, and plenty of flows are tiny. The real dividing line is about what the platform can see. A script is something Windmill watches from the outside. A flow is something Windmill watches from the inside. That difference — not size — is the whole decision.
Think of an atom and a molecule.
A script is an atom: the smallest piece of automation that stands on its
own. One nucleus of logic, one main(), a complete and stable thing. You can
hold it by itself, run it by itself, reason about it by itself. It needs
nothing bonded to it to exist.
A flow is a molecule: atoms bonded together. And here is the part that matters — the molecule is not just "more atoms." The bonds — the wiring between steps — give the molecule properties that no lone atom has. Because the atoms are bonded in a known structure, you can do things to one atom without disturbing the others: retry a single atom if it misbehaves, run two atoms at the same time, branch to a different atom depending on a result, and watch the whole reaction happen one atom at a time. A lone atom has none of those handles. It just is, or it isn't.
A script is one file with one main(). To Windmill, it runs as a single
opaque unit. The platform sees the outside of the box: it started, it finished
or it failed, here are the logs and the final result. It does not see
anything structured happening inside — no intermediate values, no per-section
timing, no place to intervene partway through.
A flow is an ordered graph of steps, and each step is usually a script. Windmill sees every step: its inputs, its output, its duration, its retries, its state. The flow is transparent from the inside; the platform can watch the data move from one step to the next and act on what it sees.
So the dividing question is never "is this job big?" It is: do you need the platform to see, or act, between the parts? Reach for a script when you have one cohesive job that either runs or doesn't — there is nothing useful to do halfway. Reach for a flow when you want per-step retry, branching, parallelism, scheduling for the whole, or simply visibility into the middle. And note the thing people miss: a flow step is a script. Same species. A flow does not introduce a new kind of building block — it bonds the ones you already have.
If size is not the test, what is? The choice comes down to a short list of yes/no questions, and every one of them is about the gaps between your steps, not the steps themselves. Do you need to retry one part without the others? To run two parts at once? To branch on a result? To watch progress? If every answer is no, an atom is enough.
Here is one cohesive job as a script — one opaque unit:
def main(env: str):
files = list_inbound(env) # read
rows = parse_and_clean(files) # transform
insert_into_datamart(rows, env) # write
return {"inserted": len(rows)}
Windmill sees this start, run, and finish. It does not see files or rows.
And here is the same work as a flow — note what the flow's modules list
actually contains:
value:
modules:
- id: a
value:
type: script
path: f/nsc/s10_list_inbound_files # <- a script
- id: b
value:
type: script
path: f/nsc/s20_process_and_load # <- another script
The flow did not replace the script. The flow is made of scripts — every
module points at one. What the flow adds is the wiring around them: the
order, and the input_transforms that hand step a's output to step b. That
wiring is the lesson in How a Flow Works; the act of cutting one
main() into those steps is From a Python Script to a Flow.
A few instincts lead people the wrong way here:
A script is one box the platform sees from the outside; a flow is boxes the platform sees between — choose by how much you need to see.
Your script needs a database password — you will not type it into the script. Here is the Windmill way.
Your script needs to talk to a database. That means it needs a host, a port, a
user, and — the awkward part — a password. The instinct from a lifetime of
scripting is to type those four things straight into the code, or into a
config.ini sitting next to it.
Don't. Windmill has a purpose-built answer for "the credentials and settings my script needs," and it is called a Resource. The same mechanism that keeps the password out of your source code is also the reason you can point one unchanged script at a test database today and the production database tomorrow — without editing a single line.
Think of electrical plugs and wall sockets.
A wall socket has a standard — a shape. How many holes, how they are arranged, what each one carries (live, neutral, ground). That standard is not a particular socket anywhere; it is the agreement on the shape. Any device built to that standard will fit any socket built to that standard.
An actual socket on the wall is a different thing. It has the standard shape, yes — but it is also physically wired to a real supply behind the plaster. Two sockets across the room can share the exact same shape and still be connected to completely different circuits.
Your script ships with a plug, cut to a socket shape. It does not carry a supply of its own. You plug it into the socket labeled "prod" or the one labeled "test" — same plug, same script — and a completely different system hums to life on the other side.
Windmill splits that idea into two named things, and the whole concept lives in keeping them straight.
A Resource Type is the socket shape — a schema. It is a named structure
with typed fields. The built-in postgresql resource type, for example,
declares that a Postgres connection has a host (string), a port (number), a
user (string), a password (string), and a dbname (string). The resource
type defines what fields exist and what type each one is. It holds no real
values; it is pure shape.
A Resource is the wired-up socket — a stored, filled-in instance of a
type. It lives at a path, like f/resources/db_prod, and it carries the actual
values: the real host, the real password. Sensitive fields are encrypted at
rest, so a resource is a safe place to keep a password in a way that a
config.ini never is.
The two click together at the script's edge. When you type a script parameter
as a resource type, Windmill recognizes it and, on the run page, shows a
resource picker instead of a blank text box — a dropdown of every resource
of that type you are allowed to use. At run time, the script does not receive a
path or a picker; it receives the resource's contents as a plain dictionary of
values. You reference resources by path, never by pasting credentials into
code. One script with a parameter typed db: postgresql can be aimed at prod
or at test simply by what you pick.
Picture one script. Out of it runs a single cable ending in a typed plug — the
parameter db: postgresql. On the wall are two real sockets of that identical
shape: f/resources/db_test and f/resources/db_prod, each wired to a
different database. The script never changes. The only decision is which socket
the plug goes into, and you make that decision at run time.
That is the whole payoff: the script depends on the shape, and the environment swap is just choosing a different socket.
Here is a script with a resource-typed parameter. Read the signature first:
# 'postgresql' is a Resource Type — the socket shape.
# Typing the parameter with it makes Windmill show a resource picker.
def main(db: postgresql):
# At run time, 'db' arrives as a plain dict of the resource's values.
host = db["host"]
port = db["port"]
user = db["user"]
password = db["password"] # decrypted by Windmill just for this run
dbname = db["dbname"]
conn = connect(host=host, port=port, user=user,
password=password, dbname=dbname)
...
You never wrote a credential. You declared a shape (postgresql) and let
whoever runs the script choose the socket. Contrast that with the
anti-pattern this replaces:
# DON'T: credentials welded into the source.
def main():
conn = connect(host="10.0.0.7", port=5432, user="svc",
password="hunter2-prod", dbname="nsc") # bad
The second version leaks a password into version control, and switching to test
means editing — and probably re-deploying — the script. The first version
switches with a dropdown. For the full before-and-after of moving a
config.ini into resources, see From config.ini to Resources.
A resource is not a secret, and a secret is not a resource. It is tempting to treat them as the same "hidden value" idea. They are not. A resource is the whole configuration object — host, port, user, password, dbname, all together. Some of its fields happen to be sensitive, and those fields are stored encrypted. The encryption is a property of certain fields, not the identity of the thing. For where standalone secrets and variables fit, see Variables vs Secrets.
You rarely invent a Resource Type. Coming from "I'll define my own
config structure," you expect to create types constantly. You won't.
Resource types are shared shapes — postgresql, mysql, s3, and dozens
more ship built in, and integrations bring their own. Day to day you create
Resources — new wired-up sockets of an existing type. Creating a brand
new resource type is the rare event, reserved for a system nothing built-in
describes.
Your script depends on the TYPE, not on any specific resource. The
parameter says db: postgresql. It does not say db_prod or db_test.
The script is bound to the shape and is deliberately ignorant of which real
database it will get. That ignorance is not sloppiness — it is exactly what
makes the test/prod swap free. A script that knew its specific resource
would be a script you had to edit to move.
It's a path, not a value. A resource is referenced by its path, and the
values live in one place behind that path. Rotate a database password? Edit
the resource once. Every script that references f/resources/db_prod
immediately uses the new password on its next run — no redeploy, no hunt
through the codebase. Inlined credentials would mean finding and editing
every copy, and hoping you missed none.
Hardcode nothing — your script carries a plug (a resource type), and the socket (the resource) decides which real system powers it.
A Windmill Variable and a Secret look like the same thing — and the difference is one checkbox that matters a lot.
In the Windmill UI, a Variable and a Secret sit in the same list, on the same page, created by the same button. They look like the same thing — and underneath, they almost are. The whole difference is a single checkbox you tick when you create one. But that checkbox decides whether the next person to open your workspace can read your database password off the screen. Getting it right matters far more than one checkbox usually does.
Think of a desk with a row of labeled drawers.
A Variable is an open drawer. It has a label on the front — default_folder,
recipient_list — and inside is a slip of paper with a value on it. Anyone
walking past the desk can pull the drawer and read what is written. That is not
a flaw; for most things it is exactly what you want. You want to glance at the
config and see what it says.
A Secret is the very same drawer — same wood, same label, same slip of paper inside — but with a small padlock on the handle. A script that holds the key can still open it and use what is inside. But the drawer never opens for reading: walk past the desk and all you see is the locked handle. In the UI, the value is shown as a row of dots.
Here is the part that surprises people: the lock is one-way. Once you lock the drawer, you cannot peek — not even you, not even as the owner. You can put new contents in, but you can never check what the old contents were. A lock that let you peek would not be a lock.
A Variable is a named value stored at a path, like f/folder/my_var. It is
workspace-scoped, so any script in the workspace can reach it; it is reusable, so
you set it once and many scripts share it; and it is plain text, fully visible in
the UI. It is the right home for non-sensitive configuration.
A Secret is a Variable with one extra thing: the "secret" flag turned on. Flip that flag and three things change. The value is encrypted at rest in Windmill's database. The value is no longer shown in the UI once you save it. And access is controlled — Windmill can track and restrict who reads it. The storage, the encryption, the hidden display: that is the lock. Everything else is identical.
The deep point is that these are not two mechanisms — they are one. A secret is a variable with a lock. This is also why Resources & Resource Types works the way it does: a resource's sensitive fields (the password in a database resource, the token in an API resource) are secret variables underneath. The resource is just a tidy bundle of plugs; the sensitive plugs are locked drawers.
So the rule for choosing is simple. Use a Variable for config that is not sensitive — a default folder path, a feature flag, a list of report recipients. Use a Secret for anything that grants access: tokens, passwords, API keys, connection strings. If leaking it would let a stranger do something, lock the drawer.
The lock is real, but it is worth seeing exactly where it sits — because that is also exactly where it stops. The encryption protects the drawer: the value as it rests in Windmill's storage and as it appears in the UI. It does not follow the value out of the drawer. The moment your script reads the secret, you are holding plain text, and the lock has nothing more to say about what you do next.
Here is the part that catches people off guard: a script reads both kinds the exact same way.
import wmill
# A plain variable — stored as plain text:
folder = wmill.get_variable("f/nsc/default_folder")
# A secret — stored encrypted, hidden in the UI:
api_key = wmill.get_variable("f/nsc/stripe_api_key")
There is no get_secret(). The secret flag is a property of the variable, not
a different API. Windmill knows the drawer at f/nsc/stripe_api_key is locked,
decrypts it for you, and hands back the value — same call, same return type. From
your code's point of view, a secret and a variable are indistinguishable. The
lock lives in storage and in the UI, not in the function you call.
1. A secret is secure storage, not secure code. The encryption protects
the value while it sits in the drawer. It does nothing about what your script
does after it opens the drawer. If your code does print(api_key) — or logs it
inside an error message, or includes it in a dict you dump — that value is now
sitting in the run logs in plain text, readable by anyone who can see the logs.
The lock is on the drawer, not on your hands after you have opened it.
2. You cannot read a secret back in the UI — by design. This is not a missing feature. A lock that let you peek is not a lock. So if you forget what a secret's value was, you do not recover it; you replace it. Keep the original somewhere trustworthy when you first set it, because Windmill will never show it to you again.
3. Don't downgrade a secret to a plain Variable "to see it easily." It is tempting when debugging: untick the box, glance at the value, move on. But the moment you untick it, the value becomes plain text — and plain text travels. It goes into workspace exports and into git syncs. A credential that was safely locked is now sitting readable in a file in a repository. Debug it some other way.
4. A credential hardcoded as a string default in your code is not a secret at
all — no matter how private the repo feels. api_key: str = "sk_live_abc123"
is plain text in your source, in every clone, in every diff, in every CI log. The
secret flag protects values stored as variables; it cannot protect a string
you baked into the code. The Default That Vanished reaches the same conclusion
from a different angle: what happens to a value is decided by where it lives,
not by how sensitive or how important it feels.
A secret is just a variable with a lock on the drawer — the lock protects the storage, not your code; never print what you unlock.
When you hit Run, your code does not run where you are sitting — and that one fact explains the most confusing bugs in Windmill.
When you press Run in Windmill, it feels like your code runs right there — same screen, same click, instant logs scrolling back at you. It does not. Your code runs on a worker: a separate process, on a separate machine, possibly in a separate building. The button is in front of you; the execution is not.
That single fact — your code does not run where you are sitting — is the quiet cause behind the most baffling bugs in Windmill. Once you really believe it, half of those bugs stop being mysterious.
Picture a restaurant kitchen with a ticket rail — that little metal strip where order slips hang in a row.
You are at the front. You write the order on a slip. You never touch a pan, never light a burner, never plate a thing. You clip the slip onto the rail and walk away. The rail is the queue. A cook — there are several of them, and you do not get to pick which one — reaches up, pulls the next slip off the rail, and cooks it. That cook is a worker.
Here is the part that matters: the cook's kitchen is not your kitchen. It has its own pantry — its own stock of ingredients, which is to say its own installed dependencies. It has its own street address — its own position on the map, which is to say its own network location. It has its own counters and cutting boards — its own filesystem, wiped clean between orders. You wrote words on a slip. Everything those words turn into happens somewhere you have never stood.
Strip away the analogy and the mechanism is plain. When you trigger a run — a script or a flow — Windmill does not execute it on the spot. It creates a job and puts that job into a queue. The queue is a shared list of work waiting to be done.
A worker is a process running on some machine in the Windmill deployment. Its whole life is a loop: pull the next job off the queue, run it, report the result, repeat. A deployment usually has many workers, and which one picks up your job is not your decision — whichever worker is free first takes it. This is exactly why Windmill scales: jobs piling up? Add more workers, and the queue drains faster. Nothing about your code changes.
That worker is its own little world. It has its own installed dependencies,
built from the dependencies your script declares in its #requirements header
(that build is its own story — see #requirements). It has its own
network vantage point — the set of hosts and databases it can actually reach,
which is decided by where the worker machine sits, not where you sit. And it
has its own ephemeral filesystem — scratch space that is not your laptop's
disk, not shared with other workers, and not kept after the job ends. Because
that environment is rebuilt the same way every time, runs are reproducible: the
same script with the same requirements behaves the same on any worker.
The clearest way to hold all this is a picture of places. Your laptop is one place on the network. The worker is a different place — different machine, different address, different set of things it can reach. The queue sits between you, a relay, not a road your data travels down.
Almost every confusing Windmill bug lives in the gap between those two places. When something "works for you" but fails when you Run it, you are not looking at a code bug — you are looking at the distance between your kitchen and the worker's.
There is no special "worker code" to write — and that is the point. The worker
is invisible in your script. What makes the worker predictable is one small
header you do write: the #requirements: block.
#requirements:
#pandas==2.2.0
#requests==2.31.0
On every single run, the worker reads that list and builds the exact
environment it describes — those versions, no others — before your main()
ever executes. It does not reuse a half-set-up environment from a previous job,
and it does not assume a package is "probably already there." It builds from
the declaration. That is the whole reason a run on Tuesday and a run on Friday,
on two different worker machines, behave identically. The worker is reproducible
because the requirements are explicit. The full mechanics of that header are
in #requirements.
Your intuition was trained on running scripts on your own machine. The worker breaks four of those instincts:
"It works on my laptop" is not "it works on the worker." They are different machines in different network positions. A database you reach through a VPN, or a host that is only visible inside one office subnet, may be flatly unreachable from the worker. The code is fine; the vantage point is not — this is the whole subject of The Code Works, the Network Doesn't.
The worker's filesystem is ephemeral and not shared. You cannot write a file in one run and read it back in the next run — a different worker may take that next job, and even the same worker has wiped its scratch space. You also cannot read it from your laptop; it was never on your laptop. Anything that must survive goes to S3 or a database, never local disk.
You do not choose the worker. Never write code that assumes a specific machine, a fixed local path, or a warm in-memory cache left over from a previous run. The next run may land on a worker that has never seen any of that. Treat every run as starting cold, on a stranger's machine.
The worker only has what the script declares. If a package is not in
#requirements, do not assume it is "probably installed" — the worker built
its environment from your declaration and nothing else. An undeclared
dependency is a missing dependency (#requirements).
You order; a worker cooks — somewhere else, in its own kitchen; 'works here' only counts if 'here' is the worker.
In a flow, step five can use the output of step one — not because the data trickled down, but because step five reached back and took it.
In a Windmill flow, step five can use the output of step one. Not step four — step one, four boxes back. Your instinct says the data must have trickled down: step one handed it to step two, step two passed it to step three, and so on until it arrived. That is not what happens. Step five reached back across the whole flow and took step one's output directly. Nothing trickled. Here is the mechanism that makes that possible.
Picture a conveyor belt running the length of a workshop, and along the belt, a row of open bins. Each bin has a label stamped on its side.
When a step finishes its work, it does not walk over and hand its output to the
next worker. It drops its output into a bin — and the bin is stamped with that
step's own id. Step a drops its output in the bin labeled results.a. Step
b drops its output in results.b. The belt never clears the old bins. Once
something is in a bin, it stays there for the rest of the run, in plain view.
So a later worker does not wait for a handoff from the worker beside it. It looks down the belt, finds the bins it needs by their labels, and reaches back to lift exactly those. A worker can grab the bin right next to it, or a bin from the very start of the line — same gesture, same belt.
And there is one bin already sitting on the belt before any worker has done a
thing: the bin labeled flow_input. That is whatever the run form collected
from the person who started the flow — the parameters they typed or picked.
It is on the belt from the first second, and any step can reach for it too.
Strip the analogy away and here is the plain mechanism. When a step finishes,
Windmill takes its return value and stores it, keyed by the step's id.
You read that stored value as results.<id> — and if the value is a dict or
list, you dig in with results.<id>.field or results.<id>.rows[0]. The
flow's own inputs — what the run form collected — are stored too, and you read
them as flow_input.<param>.
How does a step say which stored values it wants? Through its
input_transforms. For every parameter the step's code expects, the
input_transforms block holds one small expression: results.a,
flow_input.env, results.b.rows[0], or just a plain static value typed in
directly. Each expression answers one question — where does this argument come
from?
This is the part worth slowing down on. Data is not streamed from one node
to the next. The entire results store is available to every step. Each step
simply pulls by name the pieces it declared. That is why step five can use
step one's output with no help from steps two, three, or four: it just writes
results.<id of step one> and the value is right there. A step may use the
result of any earlier step, not only the one immediately before it. This is
the same wiring you met in How a Flow Works — only now you are looking at
it from the data's side of the glass.
To make this concrete, zoom in on a single step and look at its
input_transforms. Each line is one argument, and each argument names the bin
it is drawn from. One argument might come from results.a, another from
flow_input, a third from a static value. The diagram below puts one step
under the magnifying glass so you can see every argument tracing back to its
source bin.
Here is the input_transforms block for one step — step b — that needs three
arguments. Read each expr as "this argument comes from that bin":
- id: b
value:
type: script
path: f/nsc/s20_process_csv
input_transforms:
files:
type: javascript
expr: results.a # the whole output of step a
env:
type: javascript
expr: flow_input.env # a parameter from the run form
expected_count:
type: javascript
expr: results.a.count # dig into step a's output: its 'count' field
Three arguments, three sources. files takes all of step a's output.
env reaches back to the flow_input bin that was on the belt before the
flow even started. expected_count reaches into step a's bin and lifts out
just one field. None of this is a handoff from the previous step — it is step
b naming, one line at a time, exactly what it wants. The surrounding flow
structure that holds this block is X-rayed in How a Flow Works.
Data does not flow down the arrows. Every step can see every earlier
result, all at once. The arrow you see on the canvas from a to b means
one thing only: "b runs after a." It does not mean "a's output travels
along that arrow into b." Whether b actually pulls a's output is a
separate fact, recorded in b's input_transforms. The arrow is execution
order; the input_transform is data. They are two different things drawn as
if they were one.
You are not limited to the previous step. Step d can pull results.a
directly — it does not have to route through b and c. The only real
limit is time: you can read the result of any step that ran earlier,
and you cannot read a result that does not exist yet. A step reaching
forward for a result that has not been produced is the one thing the belt
cannot give you.
What goes in the bin must be serializable. A bin holds JSON-friendly
values — numbers, strings, lists, dicts. It cannot hold a live object, an
open database cursor, or a file handle. If step a tries to return one of
those, the bin gets nothing useful and step b finds it empty. This is the
same boundary rule explained in From a Python Script to a Flow: across a step
boundary, only serializable data crosses.
Rename a step's id and every reference to it breaks. The id is the
label on the bin. Change step a's id to fetch, and every
results.a expression elsewhere in the flow is now reaching for a bin that
no longer exists. Renaming an id is not cosmetic — it is relabeling a bin
that other steps are still pointing at.
A flow doesn't pass data hand to hand — every result sits in a labeled bin on the belt, and each step reaches back for the bins it names.
Two Windmill constructs are both called 'branch' — one takes a single path, the other takes them all, and picking wrong is a classic bug.
Windmill gives you two ways to split a flow, and — unhelpfully — both have
"branch" in the name: branchone and branchall. They look like siblings on
the canvas, two boxes that fan out into smaller boxes. They are not siblings.
One of them takes a single path and ignores the rest. The other takes
every path at once. Reach for the wrong one and your flow will run — it just
won't do what you meant. That confusion is one of the most common branching
bugs there is.
Picture a road that forks ahead of you.
The first kind of fork is signposted. At the split there is a stack of
signs, one per road, and each sign carries a condition: "this road if it is
raining," "this road if it is a weekday," "this road otherwise." You read the
signs from the top down, and you take the first road whose sign is true.
The moment you commit, the other roads are done — nobody travels them. One
traveler, one road. That is branchone.
The second kind of fork is different. The road simply splits into several
parallel lanes, with no signs and no choosing. A copy of you drives every
lane — all of them, at the same time. Further down, the lanes merge back into
one road, and each copy of you hands in a report from the lane it drove. You
end up with a stack of reports, one per lane. That is branchall: not a
decision, a fan-out.
So the question to ask is never "which branch construct" — it is "am I
choosing a path, or doing all of them?" Choosing is branchone. Doing all
of them is branchall.
Strip the analogy away and look at the mechanism.
A branchone is a list of branches, and each branch carries a predicate
— a small JavaScript expression that evaluates to true or false. When the flow
reaches the branchone, Windmill walks the predicates in order and runs
the first branch whose predicate is true. The rest are skipped. There is
also a default branch for the case where no predicate matched. Exactly one
branch runs — always one, never zero, never two. If you know Python, this is
plainly if / elif / else: a router that picks one path.
A branchall has no predicates at all. It is a list of branches, and
every branch runs — by default in parallel, the way the parallel lanes ran
together in Parallel Steps. The flow waits at the merge point until all
branches have finished, then collects their outputs into a list: branch 0's
result, branch 1's result, branch 2's result, in order. That is not a decision
— it is a fan-out.
The one-line rule: branchone is for routing — which path do I take.
branchall is for fan-out — do many things, then gather the results.
The output of a branchall deserves a close look, because it is where people
get surprised. The result is a positional list: an array indexed by the
branch's position in the list — 0, 1, 2 — not by any name or label you gave
the branch. Downstream steps read a branch's output as results.<id>[0],
results.<id>[1], and so on. The diagram below shows three branches feeding
their outputs into one indexed list — and what happens to those indexes when
you drag a branch to a new spot.
Here are both constructs side by side, as they appear inside the modules
list you met in How a Flow Works.
A branchone — branches each carry a predicate, plus a default:
- id: route
value:
type: branchone
branches:
- summary: weekend run
expr: flow_input.day == "saturday" || flow_input.day == "sunday"
modules:
- id: a
value: { type: script, path: f/nsc/weekend_report }
- summary: month-end run
expr: flow_input.is_month_end == true
modules:
- id: b
value: { type: script, path: f/nsc/month_end_report }
default: # runs when NO expr above was true
- id: c
value: { type: script, path: f/nsc/daily_report }
A branchall — branches, no predicates, every one runs:
- id: fanout
value:
type: branchall
branches:
- summary: notify slack
modules:
- id: d
value: { type: script, path: f/nsc/post_to_slack }
- summary: write the archive
modules:
- id: e
value: { type: script, path: f/nsc/write_archive }
In the branchone, exactly one of a, b, c runs. In the branchall,
both d and e run, and results.fanout becomes [d's output, e's output].
branchone takes the FIRST true predicate, not the "best" one.
Windmill does not weigh the branches and pick the most specific match — it
stops at the first expr that returns true and never looks further. So
order is logic. Put your specific conditions above your general ones. A
broad predicate sitting first will swallow every run that a narrower
predicate below it was meant to catch.
branchall's result is POSITIONAL. The output is a list indexed by
branch position — [0], [1], [2] — not by the branch's summary or any
name. Reorder the branches in the editor and every index quietly shifts
under you: the step that read results.fanout[1] is now reading a different
branch's output, with no error to warn you. As How Steps Pass Data
shows, downstream steps pull by exact reference — and the reference here is
a number that moves.
branchall runs ALL branches — there is no skipping. There is no
predicate, no condition, no "run this one only if." If a branch is in the
list, it runs. A branch that should sometimes not run does not belong in a
branchall at all — gate it with a branchone instead, or move the
condition inside the branch.
A branchone with no match and no default runs nothing. If none of the
predicates is true and you never defined a default branch, the construct
simply does nothing and the flow moves on. No error, no warning — just a
silent gap where you expected work to happen. Always provide a default
unless "do nothing" is genuinely a valid outcome.
branchone takes one road by condition; branchall takes every road at once — route with one, fan out with the other.
You have 200 files to process — you do not write the processing 200 times, and you describe the loop just once.
You have 200 files to process. You do not draw 200 boxes on the flow canvas,
one per file. You do not even reach for a Python for and bury the work inside
a single step. You describe the processing once — and you hand Windmill the
list of 200. The flow does the repeating. The number 200 could be 2 or 20,000
and your flow would not change a single line.
Think of a rubber stamp.
You carve the stamp once. The carving is slow, careful work — you cut the design into the rubber, check it, fix it. But once it is carved, the stamp is done. It does not change again.
Then you take a stack of blank papers and you press. Press, slide the next paper under, press, slide, press. Every sheet comes out with the same mark, because it is the same stamp. The thing that changes between presses is not the stamp — it is the paper underneath it.
A forloop is exactly this. The carving is the loop body: the steps you put inside the loop, designed once and frozen. The stack of papers is your list of items. Each press is one iteration. You never re-carve the stamp for sheet number 147; you just slide sheet 147 into place and press. The design is fixed; the item under it changes every time.
Strip the analogy away. A forloop is a flow module — one node in the
modules list you met in How a Flow Works — but instead of running one
script, it has two parts:
results.a (the output of an earlier step), or flow_input.files, or a
literal list.modules list — the steps that run inside the
loop. This is the carved stamp.Windmill takes the iterator's list and, for each item in it, runs the entire
body once. Inside the body, the current item is on the belt for you to grab:
flow_input.iter.value is the item itself, and flow_input.iter.index is its
position in the list (0, 1, 2, ...). So the body is identical every iteration,
but iter.value is different every iteration — the paper under the stamp.
Two things make this more than a convenience. First, the iterations can run
in parallel — by default Windmill runs several at once, capped by a
parallelism limit — or you can force them strictly sequential. Second, the
forloop's result is a list: one entry per item, in iterator order,
collecting each iteration's output. Press 200 times, get a stack of 200 marked
sheets back.
And here is the part that matters most: this is not a Python for buried
inside your code. It is a loop at the flow level, which means Windmill sees
every single iteration. Each press is its own job — visible on the run page,
individually retriable, individually inspectable. The repetition is not hidden
inside a step; it is the structure.
There are two ways to process 200 files, and they look almost the same from a
distance — but they are profoundly different up close. One is a flow-level
forloop: 200 iterations Windmill can see, retry, and inspect one by one. The
other is a single script step with a Python for inside it: one opaque job
where the repetition happens behind closed doors. The diagram below puts them
side by side so the trade is obvious.
Here is a forloop module in flow YAML. Step a produced a list of files;
the forloop iterates over it and runs one script per file:
- id: b
value:
type: forloop
iterator: # the stack of papers
type: javascript
expr: results.a # a list, evaluated once when the loop starts
parallel: true # run iterations concurrently...
parallelism: 4 # ...but no more than 4 at a time
modules: # the carved stamp — the loop body
- id: c
value:
type: script
path: f/nsc/s20_process_one_file
input_transforms:
file:
type: javascript
expr: flow_input.iter.value # THIS iteration's item
position:
type: javascript
expr: flow_input.iter.index # THIS iteration's index: 0, 1, 2...
The iterator is the list. The inner modules is the body. And inside the
body, flow_input.iter.value is how a step grabs the item it is working on
this press. The surrounding flow structure that holds this module is X-rayed in
How a Flow Works.
A flow-level forloop is not a Python for. They feel
interchangeable; they are not. Use the flow forloop when you want each item
to be visible — its own job, its own logs, its own retry. Use a plain
Python for inside one step when the items are trivial and you do not care
to see them individually. The danger is the wrong scale: a forloop over
10,000 tiny items is 10,000 jobs, with 10,000 jobs' worth of scheduling
overhead. If the work per item is one cheap line, loop in Python; if it is a
real unit of work, loop in the flow.
Iterations are parallel by default. They do not run one-after-another
unless you ask. That means they do not share state — iteration 5 cannot
see a variable iteration 4 set — and you must not assume order:
iteration 9 may well finish before iteration 2. (The final result list is
still in iterator order; the running is not.) And watch the parallelism
limit against rate-limited APIs — 50 parallel iterations hammering an API
that allows 5 calls a second will get you throttled or banned.
The iterator is evaluated once, up front. The list must already exist when the loop starts. You cannot grow it mid-loop, and an iteration cannot add an item for a later iteration to pick up. Windmill reads the iterator expression a single time, freezes the list, and presses the stamp that many times — no more, no less.
The body is the same for every item. That is the whole point of a
stamp. If iteration 3 needs to behave differently from iteration 4, you do
not edit the steps — the steps are frozen. You vary behaviour through the
item's own data (iter.value carries whatever distinguishes it) or with
a branch inside the body, the kind covered in
branchone vs branchall. Want some iterations to truly run at once and
rejoin? That is the territory of Parallel Steps. The loop body itself
never changes shape.
A forloop is one stamp pressed onto many items — design the body once, and Windmill runs it (and sees it) per item.
Two constructs decide 'should the flow keep going?' — one loops the body again, one stops everything; reason about them backwards and they bite.
Two Windmill constructs answer the same question — should the flow keep
going? — and they answer it in opposite directions. A whileloop says "keep
going with the body: run it again." A stop_after_if says "do not keep going
at all: end here." Both are just conditions. And because both are conditions,
both bite the moment you read them backwards.
Picture a relay race on a track.
At the start-finish line stands a lap check. Every time a runner crosses
it, the check asks one question: another lap? While the answer is yes, the
runner turns and does the loop again — same stretch of track, one more time.
The check does not know in advance how many laps there will be. It just keeps
asking, lap after lap, until the answer is finally no. That is a whileloop.
Further along the track stands a race marshal holding a red flag. The
marshal is not asking about laps. After a runner passes that spot, the marshal
glances at the flag: if it is raised, the race is over — right there. The
runners who would have come next simply stand down. The marshal does not undo
the lap that just finished; it only decides that nothing after it happens. That
is stop_after_if.
One gate sends you around again. One flag ends the whole race. Same track, two very different decisions.
A whileloop is a loop with no fixed list. Where a forloop is
handed a list up front and runs its body once per item, a whileloop is handed
nothing to count. It runs its body, evaluates a condition, and — while that
condition holds true — runs the body again. You do not know up front how many
iterations you will get. The number is discovered as you go. The catch built
into that design: the body itself must move toward the condition — advance a
page number, drain a queue, flip a flag — or the loop never ends.
A stop_after_if is not a step at all — it is a property you attach to a
step. It carries one expression, and Windmill evaluates that expression
after the step finishes. If the expression is true, the flow stops cleanly
right there: later steps do not run. The crucial word is cleanly. This is a
successful early exit, not an error. The flow's status is "completed," the
step that triggered it kept its result, and everything downstream is simply
skipped — the way the relay marshal ends the race without disqualifying anyone.
So: whileloop means "do the body again and again." stop_after_if means "we
are done — end here." One repeats; one halts.
stop_after_if is easy to misplace in your head because it reads like an
if-before-a-block in ordinary code. It is not. It hangs off the side of a
step that already ran, and it is checked the instant that step finishes. If
the expression comes back true, the rest of the flow — every module after it —
stands down. The diagram below puts the check exactly where it really sits: not
before a step, but after one.
Here are both constructs in flow YAML. First a whileloop module — note it has
a modules list of its own (the body) and a skip_failures-style condition
that decides whether to loop again:
- id: a
value:
type: whileloop # a loop with no fixed list
skip_failures: false
modules: # the BODY — runs each iteration
- id: b
value:
type: script
path: f/nsc/s30_fetch_next_page
# the loop repeats WHILE this expression is true:
# here, while step b reports there is another page
And a stop_after_if — a property hanging off a normal step:
- id: c
value:
type: script
path: f/nsc/s40_list_inbound_files
stop_after_if: # checked AFTER step c finishes
expr: results.c.length == 0 # no files? end the flow — cleanly
skip_if_stopped: true
The whileloop is a container with a body; the stop_after_if is two lines
bolted onto an ordinary step. Both live inside the same modules list you saw
X-rayed in How a Flow Works.
A whileloop with a condition that never flips is an infinite loop — and it costs real time and real jobs. This is not a hung browser tab; it is a loop running on a worker, spawning iteration after iteration forever. Always give it a guaranteed exit: a max-iterations guard, a hard page limit, a counter that the body increments no matter what. "It should stop eventually" is not an exit. A number you can prove is.
stop_after_if stops the flow successfully — your failure logic will not
catch it. A stopped flow is reported as completed, not failed. So a
failure_module (see retry & failure_module) never fires, and any "did it
fail?" check downstream sees green. If "this should not happen" is what you
mean, do not stop — raise an error. stop_after_if is for "we are done,"
never for "something went wrong."
stop_after_if is checked AFTER the step runs, not before. It cannot
prevent its own step. The step always executes, always produces its result;
the condition only decides whether anything after it gets to run. If you
wanted to skip the step itself, stop_after_if is the wrong tool — you want
a branch or a per-step skip condition.
Do not reach for a whileloop when a forloop fits. If you already have
the list in hand — rows to process, files to upload, IDs to fetch — that is
a forloop: clearer to read, and parallelizable across iterations. A
whileloop cannot parallelize, because it does not know iteration n+1
exists until iteration n finishes. Use whileloop for "until," and only
for "until." For "for each," use the loop built for it.
A whileloop repeats the body until a condition; stop_after_if ends the whole flow after a step — one loops, one halts.
Steps fail — a network blips, an API rate-limits — and Windmill gives you two different answers to 'what now?'.
Steps fail. Not because your code is bad — because the world is unreliable. A network blips mid-request. An API hands back a 429 because you called it once too often this minute. A database hiccups during a failover. These things are not bugs; they are weather.
So Windmill asks you a question every flow eventually has to answer: what
now? And it gives you two different answers — retry and failure_module —
which operate at two different levels. One works on a single step. The other
works on the whole flow. Knowing which is which is the difference between a
flow that shrugs off bad weather and one that falls over in a light drizzle.
Picture a parcel delivery.
A courier arrives at your door and knocks. No answer. Now — most of the time, nobody answered because you were in the back room, or the dog drowned out the first knock. The delivery did not fail; it just needs another try. So the courier knocks again. And maybe a third time. That second and third knock is a retry: a small, cheap, optimistic assumption that the problem was a passing one, and trying again will simply work.
But suppose the courier knocks three times and there is genuinely no one home, no safe place to leave it, nobody next door. Now the delivery has truly failed. That is not the courier's problem to keep solving at your door — it is the depot's. Back at the depot there is a standing procedure for a failed delivery: route the parcel to the return tray, log the attempt, send the sender a notification. That procedure is the failure_module.
One rides out the blip. The other handles the real defeat. The courier never files paperwork for a knock that just needed repeating, and the depot never sends "delivery failed" emails for a parcel still being knocked on. Each works at its own level, and that is exactly the split Windmill draws.
retry is a property on a step. You attach it to one module, and it says:
if this step fails, run it again — up to N times — with a delay between
attempts. That delay is often not constant; it grows, doubling each time. That
is exponential backoff: wait 1 second, then 2, then 4, giving an
overwhelmed API or a flapping network room to recover instead of being
hammered. If any attempt succeeds, the flow simply continues as if nothing
happened — downstream steps never even know there was a stumble. retry is
per-step, and it exists for transient failures: the blip, the timeout,
the rate-limit. It is the courier knocking again.
failure_module is a special module on the FLOW. It is not attached to any
one step; it belongs to the flow as a whole, and it runs only if the flow
fails — meaning a step failed and, if it had a retry, exhausted every
attempt. When it runs, it receives the error as its input: what broke, and
where. You use it to respond to defeat — send an alert to Slack, email the
team, write a row to an audit table, clean up a half-finished job. If you have
written Python, failure_module is the flow's except block: the one place
that catches whatever went wrong, anywhere upstream. retry says "try again";
failure_module says "the flow lost — now do the cleanup."
The two are not rivals; they are stages. A step runs, and if it fails its
retry takes over: attempt two, attempt three, each after a growing pause.
Only when that step has exhausted its retries — used every attempt and still
failed — does the flow itself count as failed. And only then does the
failure_module run.
So the order is fixed and worth memorizing: the step's own retries come first,
the flow's failure_module comes last. The failure_module is downstream of
every retry on every step. It never fires on a first stumble — it fires when
the flow has genuinely run out of road.
Here is a step carrying a retry block, and the flow's failure_module
sitting beside the modules list. If the YAML shape looks unfamiliar, the full
anatomy is in How a Flow Works.
value:
modules:
- id: a
value:
type: script
path: f/nsc/s20_call_flaky_api
retry: # attached to THIS step only
constant:
attempts: 0
exponential:
attempts: 3 # try up to 3 more times
multiplier: 1
seconds: 2 # wait 2s, then 4s, then 8s
failure_module: # belongs to the FLOW, not a step
id: failure
value:
type: script
path: f/nsc/s99_notify_on_failure
input_transforms:
error:
type: javascript
expr: flow_input.error # the error that broke the flow
Step a calls a flaky API. If it fails, its retry quietly tries three more
times with a growing delay. If even the last attempt fails, the flow fails —
and failure_module runs s99_notify_on_failure, handed the error so it can
say exactly what broke.
Your instinct is to reach for retry whenever something fails, and to treat
failure_module as a fix. Both instincts are wrong in specific ways:
retry only helps TRANSIENT failures. A retry is a bet that the cause
was a passing one. If the cause is a real bug — a null pointer, a malformed
SQL query, a missing required field — then attempt two and attempt three
will fail in the exact same way. You have not added resilience; you have
added latency. Do not retry deterministic errors. Retrying a guaranteed
failure just fails three times slower.
A retried step must be safe to run again — it must be idempotent. Retry
re-runs the whole step, not "the part that failed." If your step inserts a
row into a database and then fails on the line after, a retry runs the
insert again — and three retries can leave you with three rows. Before you
add a retry, ask: if this runs twice, is the result the same as running
it once? If not, make it idempotent first. An un-idempotent retry is worse
than no retry at all.
failure_module runs after retries are exhausted — not on the first
stumble. It is the last resort, not a per-step hook. It will not fire when
step a fails its first attempt; it fires only when the whole flow has
given up. If you want something to happen every time a particular step
trips, failure_module is the wrong tool — it sees only final defeat.
failure_module is for HANDLING failure, not FIXING it. It can notify,
alert, log, and clean up — but it cannot rewind the flow and resume from
where things broke. It is the except block, and the except block does
not un-throw the exception. Recovery — actually getting the work done — is
the job of retry, or of branching logic, never of failure_module. (For
the specific case of a worker that simply cannot reach a host, see
The Code Works, the Network Doesn't — and remember from The Worker that
the retry runs on a worker, possibly a different worker each attempt.)
Retry rides out the transient blip on one step; failure_module is the flow's safety net for when a step truly loses.
Some steps do the same expensive work over and over; cache_ttl lets a step remember its own answer for a while.
Some steps in a flow earn their keep once and then just repeat themselves.
A step fetches the same reference table. A step calls the same slow API and
gets back the same answer it got five minutes ago. The work is real and the
work is expensive — but the result hasn't changed. cache_ttl lets that
step remember its own answer, so the second time you ask, it doesn't lift a
finger.
Picture a desk with a telephone on it. To answer a question you have to pick up the phone, dial a slow office, wait on hold, and finally write down what they tell you. It works, but every single call costs you the same wait.
So the first time, you do the real work — you make the call — and the moment you have the answer, you write it on a yellow sticky note and press it to the desk. For the next while, anyone who asks the same question doesn't touch the phone at all. They glance at the note. Instant, free, no hold music.
The note doesn't last forever, and that's on purpose. You write a little time on the corner of it. When the note gets old — when that time runs out — you peel it off, throw it away, and the next question makes you pick up the phone once more. That stretch of time the note is good for is the TTL: time to live. Long enough to save you a hundred redundant calls; short enough that the answer on the note is still true.
cache_ttl is a property you set on a single step in a flow. Its value is a
number — a count of seconds. That's the whole knob.
Here is what Windmill does with it. When the step is about to run, Windmill asks one question: have we already run this exact step, with these exact inputs, within the last TTL seconds? If the answer is yes, it skips the step entirely and hands back the result it saved last time — a cache hit. If the answer is no, it runs the step for real, stores the result with a timestamp, and that stored result becomes the note for the next caller — a cache miss.
The crucial part is what counts as "the same question." The cache key is the
step plus its inputs. Different inputs are a different note — ask for
customer 42 and you get customer 42's saved answer, ask for customer 99 and
that's a miss until it too is cached. Same inputs inside the window get the
note; same inputs after the window expires get a fresh run. Those inputs are
the very results.<id> and flow_input values you met on the conveyor belt
in How Steps Pass Data — the cache watches exactly what the step pulled
in.
So cache_ttl is for a specific kind of step: one that is expensive and
whose answer does not change minute to minute. Reference data, slow
lookups, a heavy computation over inputs that rarely vary. If a step is cheap,
or its answer shifts constantly, the sticky note is the wrong tool.
The thing to picture is the cache key itself. It is not just "this step" — it is this step bundled together with the inputs it received. Change one input value and you have minted a brand-new key, which means a brand-new note. The diagram below pulls that bundle apart so you can see why two runs of the same step can land on a hit or a miss depending entirely on what flowed in.
cache_ttl lives right on the module, alongside id and value — the same
module anatomy X-rayed in How a Flow Works. Here it is on a step that
fetches a slow reference table:
- id: b # a step in the modules list
value:
type: script
path: f/nsc/s15_fetch_reference_table
cache_ttl: 3600 # remember the result for 1 hour
input_transforms:
region:
type: javascript
expr: flow_input.region # this input is part of the cache key
That single line, cache_ttl: 3600, says: for one hour, if this step runs
again with the same region, don't fetch anything — return the saved table.
Run the flow ten times in that hour with region set to "midwest" and the
slow fetch happens exactly once; the other nine runs read the note. Switch
region to "south" and that's a different key, so it fetches once more. The
same trick pays off beautifully inside a forloop, where the same
lookup would otherwise repeat on every single iteration.
The cache key is the INPUTS — nothing else. Windmill decides "same question" by looking only at the step and the arguments handed to it. If the step's real output secretly depends on something not in its inputs — the wall clock, a database table that someone else just edited, a random number, today's date — the cache cannot see that change. It will cheerfully hand you a stale or flat-out wrong answer because, as far as the key is concerned, you asked the identical question.
cache_ttl is not "run once." It is "run at most once per TTL
window." People read the word cache and imagine a one-time setup. It isn't.
When the window expires, the very next call is a miss and the step runs for
real again. A TTL of 300 means the step can run every five minutes forever —
just never twice within the same five.
Don't cache cheap steps. Caching has its own bookkeeping — storing the result, hashing the inputs, checking the timestamp. For a step that already finishes in milliseconds, you spend that overhead and take on staleness risk and get nothing back. Cache the expensive and stable. Leave everything else alone.
A cached step is SKIPPED — side effects and all. On a hit, the step's
code does not run. That sounds obvious until you remember a step can do
more than return a value: it might write a row to a database, send an
email, drop a file. On a cache hit, none of that happens — you get the
saved return value, but the write never repeats. cache_ttl belongs on
pure lookups, steps whose only job is to compute and return. Put it on a
step with side effects and you have silently switched those side effects
off for the length of the window.
cache_ttl is a sticky note on a step — within the window the same question gets the saved answer; just don't note down something that won't stay true.
If step B and step C don't need each other, why would they wait in line? In a flow, they don't have to.
A flow runs its steps in order — step a, then step b, then step c. That
is the safe default, and most of the time it is exactly right. But look closer
at any real flow and you will spot pairs of steps that have nothing to do with
each other. Step b fetches yesterday's sales. Step c fetches the weather.
Neither one reads the other's output. Neither one cares.
So why does c wait politely for b to finish before it even starts? It does
not have to. If step b and step c do not need each other, they can run at
the same time — and the wall-clock time you save by letting them is not a
rounding error. It is real seconds, every single run.
Picture a Formula 1 pit stop.
The car rolls in, and four crew members go to work — one at each wheel. They do not take turns. The front-left crew does not stand around watching the rear-right crew finish first. All four wheels are changed at the same time, because changing one wheel tells you nothing about changing another. The jobs are independent, so they happen in parallel.
Now here is the part to hold onto. The pit stop is not as long as one wheel plus another wheel plus another plus another. It is as long as the slowest single wheel. If three crew members finish in 2.1 seconds and one fumbles a nut and takes 2.8, the stop took 2.8 seconds — not 9.8. And the car does not leave when the first wheel is done. It leaves when all four are done. The crew splits the work into lanes, every lane runs at once, and the car rejoins the race only after the last lane finishes. That split-and-rejoin shape is exactly what parallel steps are.
In a flow, steps that do not depend on each other's output are free to run at
the same time. Windmill gives you a few ways to ask for that. You can use a
branchall — a construct that runs every one of its branches concurrently
(the sibling of branchone, covered in branchone vs branchall). You can
run a forloop with its iterations set to run in parallel — the same step
applied to many items, all at once. Or you can simply lay out a flow so that
several independent steps fan out from a common point.
Two motions matter, and they always come as a pair. Fan-out is the split: one path becomes N concurrent paths, each running on its own. Fan-in is the rejoin: the flow stops at a barrier and waits for every one of those paths to finish before it continues. Once they have all finished, the next step can reach for all of their results at once — the join collects every branch's output into one place.
This is where the payoff hides. The wall-clock cost of a parallel section is
not the sum of its lanes — it is the slowest single lane. Three branches of
2 seconds and one of 5 seconds cost you 5 seconds, not 11. And the thing that
forces steps back into single file is dependency, nothing else. If step c
needs step b's result, it genuinely has to wait for b — there is no result
to read until b is done. But if c does not touch b's output, that waiting
is pure waste. Independence is your permission to go parallel.
The clearest way to feel the win is to put the two layouts side by side and watch the clock. Sequential, four 3-second steps cost twelve seconds — each one waits for the last. Parallel, those same four steps fan out into four lanes and the section costs three seconds, because every lane runs at once and the join waits only for the slowest. The diagram below makes the rule visible: the height of a parallel section is its tallest lane, never the stack of all of them.
Here is a flow with a parallel section expressed as a branchall. Every branch
listed under it runs concurrently; the flow waits for all of them, then moves
on. For the full anatomy of a .flow.yaml file, see How a Flow Works.
summary: Build the morning report
value:
modules:
- id: a # runs first, alone
value:
type: script
path: f/report/s10_pick_date
- id: b # the parallel section
value:
type: branchall
branches: # all branches run AT ONCE
- modules:
- id: c
value:
type: script
path: f/report/s20_fetch_sales
- modules:
- id: d
value:
type: script
path: f/report/s21_fetch_weather
- id: e # runs only after b's join
value:
type: script
path: f/report/s30_assemble
input_transforms:
parts:
type: javascript
expr: results.b # the join: every branch's output
Step c (sales) and step d (weather) do not reference each other, so the
branchall runs them side by side. Step e sits after the join: it cannot
start until both branches finish, and results.b hands it the collected
output of every branch at once.
Parallelism feels like magic free speed. It is genuinely useful — but four instincts will lead you wrong:
Parallel is not free. Each parallel branch is still a real job that needs a real worker to run it. Splitting a flow into 50 branches does not conjure 50 workers — if your pool has 8 workers, 8 branches run and the other 42 sit in the queue waiting their turn. Parallelism is bounded by your worker pool (The Worker). Past that ceiling, "parallel" quietly becomes "batched."
The fan-in waits for the slowest branch. A parallel section finishes when its last branch finishes, not its first. One slow branch — a sluggish API, a heavy query — holds up the entire join, and every fast branch that finished long ago just waits at the barrier. Parallel is only ever as fast as your slowest lane, so the lane worth optimizing is that one.
Parallel branches must be truly independent. Branches do not share memory, do not see each other's partial results, and run as separate jobs on separate workers. They also must not race on the same external resource — if two branches write the same database row or the same S3 object, you have a race condition, not a speedup. Parallelize work that is genuinely separate.
Order is not guaranteed inside a parallel section. Branch 1 may finish before branch 2, or after it, or at the same time — it depends on the workers, the data, the network that day. Never write code that assumes one branch completes before another. If something truly must happen in order, that something is not parallel work — keep it sequential.
Independent steps don't need to queue — split the flow into lanes and rejoin; the cost is your slowest lane, not the sum.
Your instinct is one big script. Windmill rewards you for cutting it into small ones — the skill is knowing where to cut.
Your instinct, coming from Python, is to write one script that does
everything: read the files, clean the data, load the database, archive the
originals — one main(), top to bottom. Windmill lets you do that. But it
rewards you for cutting that script into pieces. The skill is knowing where
to cut.
A plain Python script is a one-person kitchen. One cook takes the order, preps the vegetables, works the grill, plates the dish, and washes up — start to finish, alone. It works. But if the grill catches fire halfway through, the cook starts the entire order again, salad and all.
A Windmill flow is a restaurant line. Prep station, grill station, plating station — each with its own cook, its own counter, its own tools. A dish moves down the line. Each station does one job and hands its result to the next. If the grill fails, you re-fire the grill — the prep station's work is already done and still good.
The cook did not get faster. The work got recoverable, visible, and parallelizable. That is the whole trade.
In Windmill, each step of a flow is its own script — its own file, its own
main(), its own dependencies, its own inputs and outputs. The flow is the
line: it decides the order and wires each station's output into the next
station's input (the wiring itself is How a Flow Works).
So "decomposing your main()" is a real, concrete task: take one function and
turn it into several scripts. The question is where to draw the lines. Three
good cut points:
A good cut leaves you with steps you can name, test, and retry one at a time.
Before — one script, one main(), everything tangled together:
def main(env: str):
files = list_inbound(env) # I/O: read
rows = []
for f in files:
rows += parse_and_clean(f) # transform
valid = validate(rows) # transform
insert_into_datamart(valid, env) # I/O: write
archive(files) # I/O: write
return {"inserted": len(valid)}
After — four scripts, each a step, wired by the flow:
s10_list_inbound_files main(env) -> list[str] (I/O: read)
s20_process_csv main(files) -> list[dict] (transform)
s30_validate_data main(rows) -> list[dict] (transform)
s50_insert_datamart main(valid, env) -> {"inserted"} (I/O: write)
Each main() is now small enough to read on one screen, test on its own, and
retry on its own. The flow's YAML just says "run s10, then s20 with
results.s10, then ..." — and that wiring is the lesson in
How a Flow Works.
The thing that crosses a step boundary is the return value — and only the
return value. In one Python script, every function shares the same memory:
a variable set at the top is visible at the bottom, an open file handle is
usable anywhere, a module-level constant is just there. Across flow steps,
none of that is true. Each step runs in its own process, possibly on a
different worker machine. The only bridge between step a and step b is:
a returns something, and b asks for it by name.
That kills three habits:
a cannot be used in step b. Each step opens and closes its own.And the opposite mistake: steps too small. One database query per step, eight steps to do one job — now you have eight processes starting up, eight connections opened, and a flow nobody can read. Granularity has a sweet spot, and the spot is "one thing you would want to retry as a unit."
Cut your script where you would want to restart it — each step should be one thing that can fail, and be retried, alone.
Your script does import pandas — on your laptop that just works; on the worker it works only if you said so.
Your script does import pandas. On your laptop, that line just works — pandas
is sitting in your environment because you pip install-ed it months ago and
forgot. On the worker, that same line works only if you said so. The worker
does not have your laptop's environment. It has exactly the environment your
script declared, and nothing else. Here is how you say so.
Think of sending your code on a trip.
Your script does not run where you wrote it. It travels to the worker — a different machine, a different place (The Worker), with its own pantry. You cannot walk over and grab a missing ingredient. You cannot assume the destination stocks what you need. Whatever your code needs at the destination, it has to bring.
So you write a packing list. Before your code ever arrives, the worker reads that list and packs a suitcase — exactly the items on the list, no more. A library you wrote down gets packed. A library you forgot to write down is simply left on the table at home. When your code lands and reaches for it, it is not there. The packing list is not paperwork; it is the literal difference between a dependency existing at runtime and not existing.
A Windmill Python script declares its third-party dependencies — and you have two ways to do it.
By default, Windmill auto-detects them. When you save a script, it scans
your import statements, figures out which packages they need, and resolves
them for you. For everyday code this is invisible and pleasant: you write
import requests, Windmill notices, and the worker has requests.
The second way is explicit: a #requirements: comment block at the very
top of the script — one package per line, each with a version. When that block
is present, Windmill uses it as the source of truth instead of guessing from
imports. You reach for it when you want the environment reproducible and
unambiguous: pinned versions, no surprises, written down where a human can
read it.
Either way, the mechanism downstream is the same. At deploy time Windmill takes
your dependencies — auto-detected or declared — and resolves them into a
locked set: every package, every transitive dependency, every exact version,
frozen. The worker builds precisely that environment once and caches it, so
every run of the script gets the identical environment. One last thing: the
standard library needs no declaration. os, json, datetime ship with
Python itself — they are already in the pantry. Only third-party packages go on
the list.
There is a quiet trap in auto-detection, and it is worth a picture. Windmill reads the name in your import statement and maps it to a package to install. Most of the time the import name and the package name are the same word — but not always. When they diverge, the mapping can guess wrong, and the only fix is to declare the real package name yourself.
Here is a small script with an explicit #requirements: block. Read the top
three lines as the packing list, and the imports below as the code that relies
on it:
#requirements:
#pandas==2.2.2
#requests==2.32.3
import pandas as pd # provided by the package 'pandas'
import requests # provided by the package 'requests'
import json # standard library — NOT declared, not needed
def main(url: str):
resp = requests.get(url) # needs 'requests'
df = pd.DataFrame(resp.json()) # needs 'pandas'
return json.loads(df.to_json()) # 'json' is free — it ships with Python
The block is plain comments — #requirements: on its own line, then one
#package==version per line. Windmill reads those comments, resolves them, and
the worker builds that exact environment before main() ever runs.
Auto-detection is convenient, not magic. It maps import names to
package names — and the two do not always match. import cv2 is the package
opencv-python. import yaml is the package pyyaml. import PIL is
pillow. When the names differ, auto-detection can resolve the wrong thing
or nothing at all, and you must pin the real package explicitly in
#requirements:.
"It is installed on my laptop" means nothing. The worker is a different
machine. It has only what your script declared and Windmill resolved. Your
laptop's globally installed packages — the ones you forgot you pip
install-ed — do not travel. If it is not on the list, it is not there.
Pin versions for anything that matters. An unpinned dependency is a
moving target: the next time the environment is re-resolved, it can quietly
jump to a newer version, and your code can break with no commit on your
side. Reproducibility does not come from hoping — it comes from pinning. A
bare pandas is a wish; pandas==2.2.2 is a guarantee.
The standard library is free. Do not list os, json, datetime,
pathlib, re — they ship with Python and are always present. Putting them
in #requirements: at best does nothing and at worst makes the resolver
stumble. Declare third-party packages only.
A related gotcha lives one step deeper: the packages your packages depend on —
the transitive deps you never typed an import for. That is its own story in
Transitive Dependencies. And when your dependency is your own shared code
rather than a PyPI package, the rules change again — see Shared Code.
The worker packs only what your script's packing list says — declare every third-party dependency, and pin it, because nothing else makes the trip.
Your script has HOST = 'prod-db.internal' near the top — it works, and it is also why you have three copies of the script.
Open your script. Near the top, before main(), there it is:
HOST = "prod-db.internal". Maybe a PASSWORD next to it, maybe a folder
path. It works. It has worked for months. It is also the precise reason you
have three slightly-different copies of this script — one for test, one for
staging, one for prod — and the reason a credential rotation means editing all
three. Here is the fix, and it is smaller than you think.
A hardcoded value is an answer written on the wall in permanent marker.
When you write HOST = "prod-db.internal" into the source, you have written
the answer directly onto the wall of the script. It is right there, it is
readable, and it cannot be changed without a paint job. So when you need the
same script to talk to the test database, you do the only thing the wall
allows: you photocopy the entire room — the whole script — carry the copy to a
different wall, and edit that wall too. Three environments, three walls, three
copies. They drift apart. One gets a bug fix the others never receive.
A Resource is the same answer printed on a small card that slots into a holder on the wall. The script does not own the answer; it reads whatever card is in the holder when it runs. One script, one room. Going from test to prod is not a photocopy — it is swapping the card. The wall never changes.
Config can live in exactly three places, and only one of them travels well.
One — hardcoded in the script. The value is a literal in the source. This is permanent marker. The value is now welded to the code, and "different environment" can only mean "different copy of the code."
Two — in a file next to the script. A config.ini, a settings.py, a
.env checked in beside main.py. This feels like externalizing, but the
file ships with the code — it is in the same repository, the same deploy
unit. It still needs a copy per environment, and any password in it is now in
git history forever. You moved the marker to a sticky note and stuck the note
to the same wall.
Three — in Windmill, referenced by path. The value lives in the platform,
outside the code, and the script names it instead of containing it. This is
the swappable card. Windmill gives you three shapes for it: a connection
bundle — host, port, user, password together — becomes a Resource
(Resources & Resource Types); a single sensitive value becomes a Secret; a
single plain value becomes a Variable (Variables vs Secrets). Your
script's main() then takes a typed Resource parameter or a variable path —
never a literal. The payoff: one script, environment-independent, with config
that is versioned and access-controlled in Windmill instead of scattered
across the code.
Picture the two worlds side by side. On the left, the hardcoded world: three
near-identical scripts, each with its own HOST and PASSWORD baked in,
quietly drifting apart. On the right, the Windmill world: one script, and
beside it two Resource cards — prod_db and test_db — either of which slots
into the same parameter. The diagram below shows what shrank: three copies
collapse into one.
Before — the values are welded into the source. Every environment needs its own edited copy of this file:
# permanent marker — baked into the script
HOST = "prod-db.internal"
PORT = 5432
USER = "etl_writer"
PASSWORD = "s3cr3t-in-git-forever" # now in version control
def main(table: str):
conn = connect(HOST, PORT, USER, PASSWORD)
return conn.query(f"SELECT count(*) FROM {table}")
After — main() takes a typed Resource parameter. The values arrive at run
time from whichever card Windmill hands in:
# postgresql is a Resource Type; Windmill renders a picker for it
postgresql = dict
def main(table: str, db: postgresql):
# db is a dict, filled from the selected Resource
conn = connect(db["host"], db["port"], db["user"], db["password"])
return conn.query(f"SELECT count(*) FROM {table}")
The script no longer knows or cares which database it talks to. Point the db
parameter at the prod_db Resource and it is production; point it at
test_db and it is test. Same file, same deploy, same git history — no
secret in it.
1. A config.ini next to the script is NOT externalized. This is the
trap that feels like the solution. The file is outside main(), yes — but it
is inside the repository. It still travels with the code, still needs one
copy per environment, and any secret it holds is now committed to git. A file
beside the code is not a Resource; it is just hardcoding with an extra step.
2. Hardcoding is not only credentials. A hostname is the obvious case, so
people fix the hostname and stop. But a hardcoded folder path, a hardcoded
email recipient list, a hardcoded threshold (if rows > 5000) is the exact
same welded-to-the-wall value. The test: if it differs between test and prod,
or if it might ever change, it is config — and config belongs in a card, not
on the wall.
3. Do not read worker environment variables to "externalize." Reaching
for os.environ["DB_HOST"] feels external — the value is not in the source.
But the worker's environment is not yours to rely on. Your script may run on
any worker in the pool, and that machine's environment is shared
infrastructure, not your script's private store (The Worker). Use
Windmill Variables and Resources — the platform-native store that every worker
resolves the same way.
4. Externalizing means the value is referenced by PATH. The point is not
merely "the value left the file." The point is that the script holds a
pointer — a Resource path like f/nsc/prod_db — and Windmill resolves it at
run time. Change the value once in Windmill and every script pointing at that
path picks up the new value on its next run. One edit, everywhere. That is the
whole reason to do this.
Stop writing config in permanent marker — a hardcoded value means a copy per environment; a Resource means one script and a swappable card.
Your script runs on a worker you cannot see — when something breaks at 2 a.m., the only witness is the log.
When you ran Python on your own machine, debugging was a conversation. You printed something, you watched it appear, you tweaked, you ran again. You were in the room.
In Windmill you are not in the room. Your script runs on a worker — a separate process on a separate machine (The Worker) — and you do not get to watch it work. So when a scheduled job fails at 2 a.m. and you read about it over coffee, there is exactly one witness to what happened: the log. Your whole job, between now and then, is to make that witness worth listening to.
Think of a flight recorder — the "black box" on an aircraft, which is, famously, painted bright orange.
You are not in the cockpit. The plane flies, climbs, banks, and lands without you watching a single instrument. But the recorder is. Every reading, every control input, every callout from the crew gets written down — timestamped, in order, and crucially, kept after the flight ends. The recorder does not care whether the flight was smooth or whether it crashed; it captures the same way either way.
When the flight is over, you do not interview the airplane. You open the
recorder and read the tape: at this minute this happened, then this, then this.
A clean flight gives you a boring tape. A bad flight gives you the exact moment
it went wrong. Either way, the recorder is the only honest account of a trip you
did not take. Your worker is the airplane. Your print() statements are the
crew's callouts. The job page is where you open the box.
Here is the mechanism, stripped of the metaphor. When your script runs on a
worker, everything it writes to stdout and stderr — every print(), every
traceback, every warning — is captured into that job's log. Windmill
streams that log to you live while the job runs, and then keeps it attached to
the job after the job is long over. Nothing is lost when the worker moves on.
The log is not the only thing a step produces, and this is the distinction that matters most. A step produces two separate things:
main(), for machines
to read. The result is what downstream steps consume; it travels the conveyor
belt between steps (How Steps Pass Data).These two channels never mix. A later step cannot read your log, and a human rarely wants to squint at a raw result. Every job — every script run, every step of a flow — gets its own log on its own job page, alongside a status (success or failure), timing, and the inputs it was called with. When a flow fails, Windmill points you straight at the failing step's job page, so you read that step's log. So observability in Windmill comes down to three habits: print your narration to the log, return your data as the result, and know how to read the job page.
The job page is the single place where all four of these come together: the log you narrated, the result you returned, the status that says pass or fail, and the timing that says how long each part took. It is the same page whether the job is a standalone script or one step inside a flow. Learn its layout once and every 2 a.m. investigation starts from the same map.
There is no special "logging API" to learn. The Python you already know is the
logging API — print() writes to the log, return produces the result.
def main(env: str, batch_size: int = 500):
print(f"Starting import — env={env}, batch_size={batch_size}")
rows = fetch_rows(env)
print(f"Fetched {len(rows)} rows from source") # a key count
valid = [r for r in rows if r.get("id")]
skipped = len(rows) - len(valid)
if skipped:
print(f"Skipped {skipped} rows with no id") # narrate decisions
print(f"Importing {len(valid)} rows in batches of {batch_size}")
imported = import_in_batches(valid, batch_size)
print(f"Done — {imported} rows imported")
# This is the RESULT — the structured value the next step will read.
return {"imported": imported, "skipped": skipped, "env": env}
Read the print() lines as a story: what the script is doing, and the key
counts at each turn. If this job fails on row 12,000, the log already tells you
it got past "Fetched", past "Skipped", and died mid-import — you have narrowed
the bug before you have read a single line of the traceback. The return at the
end is a different thing entirely: that dictionary is the result the conveyor
belt carries onward (How Steps Pass Data).
Coming from a single-machine Python habit, four instincts will mislead you:
The log is for humans; the result is for machines. Do not print() your
data hoping a later step will pick it up by reading the log. It will not — and
it cannot. Downstream steps read the result (results.x), never the log.
The log is narration; the result is the payload (How Steps Pass Data).
print is the witness — a silent script tells you nothing. A script that
does its work quietly and then simply fails leaves you with a traceback and no
context. Print what it is doing and the key counts as it goes. The difference
between a five-minute fix and a two-hour one is whether the 2 a.m. log is
actually readable.
Never print() a secret. A printed token, password, or connection string
lands in the log as plain text — and the log is kept, attached to that job,
readable by anyone who can open the job page. Pull secrets through Windmill's
secret mechanism and never let them touch stdout (Variables vs Secrets).
You cannot attach a debugger to a worker. There is no breakpoint, no stepping, no live inspection on a machine you cannot see. Your debugger is the log plus a test run. Design for that reality: log enough that the log alone can tell you what happened, because the log is all you will get.
The worker runs out of your sight — print() is the flight recorder; narrate to the log, return data as the result, and read the job page.
A cron line on a server runs your script on time — until the server reboots, or it silently stops and nobody notices.
You have a script that needs to run every morning at six. So you SSH into a
server, type crontab -e, add one line, and save. Done. It works — for a
while.
Then one of three things happens. The server reboots and comes back without that cron daemon healthy. Or you spin up three more servers and forget which one carries the line. Or the script quietly starts failing at 6:01 every morning and nobody finds out until someone asks where last week's report went. A Windmill Schedule fixes all three — because it does not live on a server at all.
A crontab line is a sticky note taped to the wall of one room.
It works fine while you are in that room. But only that room sees the note. If the room is repainted — the server rebooted, reimaged, replaced — the note is gone. And to find out whether the thing on the note actually happened, you have to walk back into that exact room and look around: SSH in, dig through log files, hope someone redirected the output somewhere readable. The note tells the future ("do this at 6"); it keeps no record of the past.
A Windmill Schedule is the alarm system wired into the building itself. It is set on a panel anyone can walk up to and read. It does not care which room you are standing in, or whether any single room still exists — it is part of the building. And it keeps a log: every time it went off, at what time, and whether the thing it triggered actually completed. You do not go hunting for evidence. The evidence is on the panel.
A Schedule in Windmill is a small, declarative object that binds three things together: a cron expression (when), a script or flow path (what), and a set of arguments (with which inputs). That is the whole definition. It is not a process. It is not a daemon. It is a record in the platform that says "at these times, run this, like this."
When a scheduled time arrives, the Schedule does exactly one thing: it enqueues a job. From that point on the run is an ordinary Windmill job — the same kind you get when you press Run by hand. A worker pulls it off the queue and executes it (The Worker). Because the Schedule only enqueues, it is not tied to any machine: the worker that runs your 6 a.m. job today and the worker that runs it tomorrow can be entirely different boxes, and the Schedule neither knows nor cares.
And because every scheduled run is a normal job, every scheduled run has a job page: a log, a result, a status, a duration. The Schedule itself keeps a history — a list you can open in the UI showing that it fired at 6:00, 6:00, 6:00, and whether each run succeeded or failed. You can see the next run time without doing any arithmetic. You can enable or disable it with a toggle, change its cron or its arguments, and the whole definition is versioned along with the rest of the workspace. The "did it run?" question that used to mean an SSH session is now a page you bookmark.
The contrast is sharpest when you put the two side by side. A crontab line is a string of text inside one file on one server — invisible to everything outside that machine, with no memory of what it has already done. A Windmill Schedule is an object in the platform that fans out into a visible trail: fire, enqueue, job, history.
The left side knows only the future. The right side knows the future and the past — and that is the difference that matters at 6:01 a.m.
A Schedule is configured, not coded. You do not write a scheduler; you fill in three fields. Conceptually, the definition looks like this:
schedule: "0 0 6 * * *" # WHEN — the cron expression
script_path: f/nsc/s30_daily_report # WHAT — the script or flow to run
args: # WITH WHICH INPUTS
env: prod
send_email: true
timezone: America/Chicago # WHICH 6am — see "where intuition fails"
enabled: true
Decode the cron expression 0 0 6 * * * field by field: seconds 0,
minutes 0, hours 6, day-of-month any, month any,
day-of-week any. Read together: "at second 0 of minute 0 of hour 6, every
day." Windmill's cron has a leading seconds field, so it is six fields, not
the classic five — a small thing that trips up people copying lines straight
from an old crontab.
The args block is the real upgrade over a crontab line. In a crontab you
bake arguments into the command string and they vanish into the file. Here
they are named, structured fields — the same inputs the script's form would
ask for if you ran it by hand.
Your intuition for crontab was built on "a line that runs on a server." A Schedule breaks four of those instincts.
A Schedule does not "run on a server" — it enqueues a job. When the cron time arrives, the Schedule drops a job onto the queue. If every worker is already busy, that job waits in line until one frees up (The Worker). So the cron time is when the job is queued, not when it is guaranteed to start. On a quiet system the gap is milliseconds; on a saturated one it can be real. If you need punctual starts, you need enough worker capacity, not just a tighter cron.
The cron expression runs in a specific timezone. "6 a.m." is
meaningless until you say where. A Schedule has a timezone field, and
0 0 6 * * * means 6 a.m. in that zone — which drifts against UTC twice
a year when daylight saving shifts. Set the timezone deliberately. Do not
leave it to a default and assume it matches the office.
A missed or failed run is visible — but only if you look. The history records every failure faithfully. It does not phone you. A Schedule has no instinct to chase you down; a red entry sits quietly in the list until someone opens the page. If "nobody noticed" is unacceptable, wire an error handler so a failed scheduled run actively alerts you (retry & failure_module). Visibility is automatic; notification is something you opt into.
Runs can overlap. A crontab line fires on the clock with no regard for whether the previous run has finished. A Schedule is the same: if your job takes seven minutes and you scheduled it every five, the next run starts while the last is still going. Sometimes that is fine. Sometimes it means two jobs fighting over the same database row. Decide which case you are in, and if overlap is dangerous, guard against it — lengthen the interval, add a lock, or design the job to detect a run already in progress.
A crontab line is a sticky note on one machine; a Windmill Schedule is an alarm in the platform — visible, machine-independent, and it remembers whether it went off.
You wrote a great helper function — you do not paste it into every script that needs it.
You wrote a clean little db_connect() helper. It opens the database the
right way: the right host, the right timeout, the right retry. Six scripts
need it. Your Python instinct says copy it into all six — or paste it once and
forget which copy is the real one. Do not. Windmill has one place for shared
code, and one way to reach it. Here is how that works.
Picture a workshop with a wall of workbenches.
On your laptop, from . import utils is a tool sitting in your own drawer.
It is right there, at your bench, easy to grab — but only the scripts in that
folder can reach it. Open a project in another folder and the tool is gone;
you packed a second copy in that drawer too. Soon you have five drawers, five
slightly different copies of the same screwdriver, and no idea which one is
sharp.
In Windmill, a shared library is a toolbox bolted to the workshop wall. It lives at a fixed spot — a workspace path — and it is not in anyone's private drawer. Any script anywhere in the workspace walks up to that toolbox and takes what it needs. And there is exactly one toolbox. Sharpen a tool — fix a bug, change a timeout — and every script that reaches for it gets the sharper tool, instantly, without you touching them one by one.
In Windmill, every script lives at a path — something like
f/folder/script_name. That path is not a filename on a disk; it is the
script's address inside the workspace. (Folders and how scripts are
organized are their own topic — see Script vs Flow.)
A library is just a regular script you do not run directly. It has no
real main() you care about — instead it holds shared functions: a
db_connect(), a clean_term_code(), a paths module. It sits at its path
and waits to be imported.
Any other script can import that library as a module, using a path-based import that mirrors its workspace path. You write the import; Windmill reads it, finds the library script at that path, and at deploy time bundles the dependency along — so the worker that runs your script also has the library's code. You never copy anything. The shared logic exists exactly once. Change it in that one place and every importer is, by definition, already updated — there is no second copy to forget.
The thing to fix in your head: a Windmill import does not point at a file next to yours. It points at an address in the workspace. The same import string works from any script, in any folder, because it names where the library lives, not where the caller lives. The diagram below traces one import from a caller, along the workspace path, to the library it resolves to.
First the library — a script at the workspace path f/common/db. It is just
functions; nobody runs it on its own:
# script at workspace path: f/common/db
# This is a library. It is never run directly — other scripts import it.
def db_connect(env: str):
"""Open the datamart the one correct way: right host, timeout, retry."""
host = "datamart-prod" if env == "prod" else "datamart-test"
# ... open and return the connection ...
return _open(host, timeout=30, retries=3)
Now a second script that needs a connection. Instead of pasting db_connect(),
it imports the library by its workspace path:
# a normal script — its own main(), its own job
from f.common.db import db_connect # import BY WORKSPACE PATH
def main(env: str):
conn = db_connect(env) # the shared helper, one source of truth
rows = conn.query("select * from students")
return {"count": len(rows)}
The import path f.common.db is the workspace path f/common/db with the
slashes written as dots. Every script that needs a database connection writes
that same line. Fix the host or the timeout once, in f/common/db, and all of
them are fixed. (Declaring the external packages a library needs is a
separate concern — that is #requirements.)
1. A Windmill import is by workspace path, not a filesystem-relative path.
from . import utils means "the file next to me." That mental model does not
map. There is no "next to me" in a workspace — there is only where in the
workspace a script lives. The import names the library's address
(f.common.db), and it reads the same no matter which folder the caller sits
in.
2. Move or rename a shared lib and every importer breaks. That import
string is a path. Change the library's path — rename the folder, move the
script — and every from f.common.db import ... line is now pointing at
nothing. Shared code has many dependents. Treat its path like a published API:
once scripts import it, do not move it casually.
3. A change to a shared lib touches every importer. This is the whole point — and the whole danger. Fix a bug once, everyone benefits. Introduce a bug once, everyone breaks. A bug in a shared library is a bug in every script that imports it. So test shared code harder than you test a one-off script; its blast radius is the entire workspace.
4. Do not make a "lib" out of one script's private helpers. Shared code earns its place on the wall by being genuinely used by many. A helper that only one script will ever call is not shared — it is that script's own business, and it belongs inside that script. A library full of single-caller functions is just indirection with extra steps.
Don't paste a helper into every script — put shared code in one library on the workspace wall, import it by path, and fix it once for everyone.
You never build the form. You declare what you need, and Windmill draws it for you.
Coming from desktop or web programming, "let the user pick the environment" means building a form: a dropdown, a label, validation, a submit button. In Windmill you build none of it. You declare what you need, and Windmill draws the form for you. The form is a side effect of an honest function signature.
Think of a customs declaration card.
You — the developer — are the customs office. You decide what questions the card asks: Anything to declare? Arriving from where? Carrying more than $10,000? You design the card by choosing the questions, never by drawing the boxes.
The end user is the traveler. They never see your decisions about layout. They see blank fields and fill them in.
In Windmill, your function's parameters are the questions on the card. Choose the parameters well and the traveler gets a clear, short, hard-to- mess-up form. Choose them carelessly and they get a confusing card — and fill it in wrong.
The run page an end user sees is generated from two things, and you control both:
main() signature. Every parameter becomes a form field. The
parameter's type chooses the widget. Its default value pre-fills the
field. A parameter with no default is required.schema Windmill infers from that signature — and that you can
refine. It carries the extras: a field's description, its allowed values,
its order. This is the schema block you meet in How a Flow Works.The type-to-widget mapping is the core of it:
You never wrote a <select>. You wrote Literal["test", "prod"], and
Windmill knew that means "exactly one of these two" and drew a dropdown.
Read a signature and you can already picture the form it will produce.
This signature:
from typing import Literal
def main(
env: Literal["test", "prod"], # required dropdown — no default
dry_run: bool = True, # checkbox, pre-checked
max_files: int = 50, # number box, pre-filled with 50
):
...
produces, with zero extra work, a run page with: a dropdown the user must
set, a checkbox already ticked, and a number field showing 50. Add
descriptions to the schema and it gets friendlier still — Windmill shows them
as helper text under each field.
The form is exactly as good as your signature — no better.
This is where it bites:
str is a free-text box. env: str lets the user type prdo,
PROD with a trailing space, production, or nothing at all. Every typo
is now a runtime bug. The fix is a type that constrains: Literal[...]
or an enum. If there is a known set of valid answers, never use plain
str.dict or list[dict] renders as a text area expecting hand-written JSON.
Fine for a developer; cruel for the office user the flow was built for.
Keep user-facing parameters flat and primitive.dry_run: bool = True opens the form pre-set to the harmless
choice. dry_run: bool = False means one distracted click runs the real
thing. Defaults are UX, not just code.max_files field is a
mystery to the user. The signature is where you would write a comment
anyway — write it as a description instead, and the user sees it too.A good Windmill form is not designed in a UI builder. It is designed in the function signature, one honest parameter at a time.
You never build the form in Windmill — you write an honest function signature, and the form builds itself; a sloppy signature is a sloppy form.
Windmill draws the form for you — the default form is correct, and often a little raw; shaping it is what makes it usable.
You already know Windmill draws the form for you — write an honest function signature and a run page appears, no UI builder required. That is the lesson of Inputs Become Forms. But there is a second lesson hiding behind it: the form Windmill draws by default is correct, and often a little raw. Every field is there, every type is right — and the office user still squints at it. Shaping the form is the small, deliberate work that turns correct into usable.
Think of a suit off the rack.
An off-the-rack suit is a real suit. It has two sleeves, the right number of buttons, a jacket and trousers that match. Technically, it fits — you can walk out of the store wearing it. But the cuffs are a touch long, the waist is loose, the shoulders sit a little wide. It fits a person; it does not fit you.
The auto-generated Windmill form is that off-the-rack suit. Every field is present and the right kind. What it has not had is a fitting. Shaping the form is the tailoring — take it in here, add a real label there, pin the hem — so it fits the actual person who will wear it: the office user who opens your flow on a Tuesday morning and just wants to launch it without guessing.
The form is rendered from the flow's JSON schema — the same schema block
you met in How a Flow Works. Shaping the form means refining that
schema. There is no separate form designer; there is just the schema, and
four kinds of refinement do most of the work:
Beyond those four, you can also constrain a value (a numeric min/max, a regex
pattern for text), and order or group the fields so related questions sit
together. And you have two places to do all of it: in the script, through
richer type annotations, or directly in Windmill's schema editor on the flow.
The annotations are version-controlled and travel with the code; the editor is
faster for a quick description. Both write to the same schema.
Four controls carry almost all the weight, and it helps to see them side by side: the enum that becomes a dropdown, the default that pre-fills, the description that becomes helper text, and the constraint that rejects a bad value before submit. None of them is a new feature to learn — each is one extra key in the schema. The diagram below lines them up so you can see the raw field on the left and what each control does to it on the right.
Most shaping happens without ever opening the schema editor — you just write a richer signature, and Windmill folds it into the schema for you:
from typing import Literal
def main(
env: Literal["test", "prod"] = "test", # dropdown, pre-set to the safe value
max_files: int = 50, # number box, sensible default
report_email: str = "ops@example.edu", # pre-filled, editable
):
"""Process inbound CSV files.
env: Which environment to run against.
max_files: How many files to process at most.
report_email: Where to send the run summary.
"""
...
The Literal becomes an enum — a dropdown — and "test" pre-selects the harmless
option. Each default pre-fills its field and makes it optional. And the docstring
lines are not just for developers reading the code: Windmill lifts them into the
schema as field descriptions, so the user sees them as helper text under each
field. One honest, well-annotated signature is a shaped form.
The defaults that feel like polish are the ones that quietly cause harm.
max_files is a fine variable name and a useless label. "How many
files to process at most" is the same idea said to the person filling in the
field. Write the description for them, not for you.dry_run = False waiting
there for a distracted Tuesday. Choose the default as if it will be accepted
unread, because it will be.main()
check the value and raise if it is wrong. But by the time your code runs,
the user has already submitted the form and waited for the worker to pick it
up — and now they get an error. An enum or a min/max stops the bad value at
the form, before submit, while the user can still fix it in two seconds.
Validate where the user is, not where your code is.The same instinct serves you in Approval Steps, where the "form" a human sees is the pause itself: a clear question beats a clever one.
The auto-form fits off-the-rack; shaping it — enums, defaults, descriptions — is the tailoring that makes it fit the person who runs the flow.
Some decisions a machine should never make alone — an approval step puts a human in the loop, on purpose.
Most steps in a flow should run the moment their turn comes — list the files, parse the CSV, write the report. But a few steps carry weight a machine should never lift alone: delete the records, send the payment, email all 4,000 students. An approval step is how you make a flow stop and ask first. It puts a human in the loop — not by accident, not because something broke, but on purpose.
Think of a drawbridge over a moat.
The flow is a vehicle driving along its route. Most of the road is open: it drives straight through. But at one point the road runs up to a raised drawbridge and simply ends. The vehicle cannot jump the gap and would not want to. So it stops at the edge and waits.
On the far bank stands a person at a lever. They look at what is waiting to cross, and they decide. Pull the lever one way and the bridge lowers — the flow drives on, across, and continues its route. Refuse, and the bridge stays up; the vehicle turns back. The wait itself costs nothing. The vehicle is not revving its engine or burning fuel idling — it is simply parked, patient, whether the person answers in two minutes or tomorrow morning.
An approval step — Windmill also calls it a suspend — is a flow module that pauses the flow and emits a request for a human decision. When the flow reaches it, two things happen. First, the flow's execution stops at that module. Second, Windmill produces a way for a person to respond: typically a link or a notification carrying an Approve and a Reject action, delivered wherever you route it — email, Slack, a Windmill page.
While it waits, the flow's entire state is held safely on the side. It remembers every step's output, exactly where it stopped, and what it was about to do next. Crucially, no worker is busy during this wait. The flow is not running a loop, not checking a clock — it is genuinely suspended, like the parked vehicle.
When a person clicks Approve, the flow resumes from that module and drives on. When they click Reject, the flow stops — it does not continue past the gate. You decide who receives the request, and you can attach a timeout so a forgotten approval does not leave the flow waiting forever.
It is tempting to picture "wait for a human" as a loop that keeps asking "are we good yet? are we good yet?" — that is how you would write it by hand. An approval step is not that. It is a true pause: the flow's state is frozen and set aside, consuming nothing, until an external answer arrives. The diagram below puts the two side by side — a paused flow holding its state with zero workers, against a polling loop that ties up a worker the whole time just to ask the same question.
An approval step is not imperative code you write — it is configuration on a flow module. You take a normal step and turn on its suspend setting, then choose who approves and how long the flow may wait. Conceptually, the module carries a small block like this:
- id: c # the step that needs a yes
value:
type: script
path: f/nsc/s40_send_emails
suspend:
required_events: 1 # how many approvals are needed
timeout: 86400 # seconds to wait — here, one day
# who gets the request, and what they see, is configured here:
# a recipient + a form summarising what they are approving
When the flow reaches step c, it pauses before running it and sends out
the approval request. One approval (required_events: 1) lets it through. If
nobody responds within timeout, the wait ends on its own — and your flow
decides what that means. The send-emails step itself stays exactly as it was;
you have only put a gate in front of it.
The wait feels expensive. It is not — and three more things surprise people.
The flow is PAUSED, not running. A suspended flow consumes no worker while it waits. It can sit at an approval step for minutes, hours, or a full day, and that costs nothing — there is no process spinning, no resource held. This is the opposite of a whileloop & stop_after_if whileloop that polls for the same answer: a polling loop occupies a worker the entire time, just to keep asking. "Wait for a human" is cheaper as a pause than as a loop, not more expensive.
An approval is a branch point, not a speed bump. "Approved" is only one of the outcomes. The person can reject, and they can simply never answer until the timeout fires. If your flow only handles the yes, a no or a silence becomes an unhandled surprise. Treat the approval like a fork in the road and decide, on purpose, what happens down each path.
Put the approval BEFORE the irreversible step, not after. The gate must guard the thing. An approval placed after the records are deleted is theater — the damage is already done and the human is only signing a receipt. Insert the suspend immediately ahead of the step you cannot take back, so the yes is what unlocks it.
The approver needs context, or the yes means nothing. A request that just says "Approve?" with no detail gets rubber-stamped — the person clicks yes because there is nothing to weigh. Carry the facts into the request: the row count, a short summary, the environment, who is affected. An approver who can see "this will email 4,127 students in PROD" is making a real decision. One who sees only a button is not.
An approval step is a drawbridge — the flow waits, free and patient, until a human lowers it; put it before the irreversible step and handle the no.
A flow gives the user a form and a Run button — sometimes that is exactly right, and sometimes the user needs a screen.
A Windmill flow already comes with a user interface. You write a function, and Windmill draws a run page: a form with one field per input, and a Run button at the bottom. The user fills it in, clicks once, and waits for the result.
Sometimes that is exactly the right amount of interface. One job, one button — ask for nothing more. But sometimes the user does not have one job. They need to look at data, pick a row out of it, run an action on that row, and see what happened — all on the same screen, without ever leaving it. When that is the need, a form with a Run button is not enough. The user needs a screen. That screen is an App.
Think of the difference between a vending machine and a control room.
A vending machine does one job beautifully. There is one panel: you make a selection, you press the button, you collect the result from the tray. You never wonder what to do next, because there is only one thing to do. It is focused, obvious, and fast. A flow with its auto-generated form is a vending machine — fill in the inputs, press Run, collect the output.
A control room is a different animal. It is a wall of buttons, dials, live screens, and a table of readings you can lean in and inspect. An operator stands in front of it, watching one display, reacting to it, reaching for the right control, then watching another display change. There is no single "the button" — there are many, and which one matters depends on what the screens are showing right now. An App is a control room: many controls and live displays, several actions, all on one screen.
The mistake is not building a control room. The mistake is building a control room when a vending machine was all anyone needed — and, just as often, bolting a fifth and sixth function onto a vending machine when the job had quietly grown into a control room.
A flow is orchestrated backend logic. It is steps, branching, loops, retries — the engine that does the work (Script vs Flow). Its user interface is not something you design; it is the auto-generated run form, inferred from your inputs (Inputs Become Forms). One form, one run. The flow's whole relationship with the user is: here are the questions, fill them in, press Run once.
An App is a Windmill front-end that you do design — by drag-and-drop. You open a blank canvas and place components on it: buttons, text inputs, tables, charts, dropdowns, labels, containers. You arrange them where you want them. Each interactive component can be wired to trigger a script or a flow, and each display component can show the result of one. The App is the cockpit; the flows and scripts are what its buttons and tables actually call.
That is the real distinction. A flow answers "fill in some inputs and run one job." An App answers "show me the data, let me click a row, then run an action on that row, and show me the outcome." An App is interactive, multi-step, and stateful on the screen — it remembers what is selected and what is loaded while the user works, in a way a one-shot run form never does.
The clearest way to picture an App is as a layer that sits on top of your existing flows and scripts, not as a replacement for them. The components on the App canvas are the visible surface; underneath, each one is a wire running down to a script or a flow that does the real work and hands a result back up.
Read the diagram top to bottom: the user touches a component, the component runs a script or flow, the result comes back and lands in a display component. The App is the UI layer — and only the UI layer.
This beat has no code listing, and that is the point: an App holds no business logic of its own. You do not write an App the way you write a script. You compose it, by placing components on a canvas and wiring each one to a runnable.
Picture a small "process inbound files" App. On the canvas you might place:
Notice that none of those components contain the logic for listing or processing files. The listing logic lives in a script. The processing logic lives in a flow. The App only decides what is on screen, what is wired to what, and where the results are shown. If you find yourself wanting to write real logic inside the App, that is the signal to move it down into a script the App calls instead.
An App is for screens, not for showing off — and it is a different kind of thing from a flow, not a bigger one. Four places intuition leads you astray:
An App is not a fancier flow — it is a different layer entirely. It is tempting to think "App = flow with a nicer UI." It is not. The App is the UI layer, and it still calls flows and scripts for the actual work. Do not put business logic in the App. Put it in the scripts and flows the App calls, and let the App do nothing but display and trigger.
Do not build an App when a flow's auto-form would do. If the whole interaction is "fill in three fields and run one job," the auto-generated run page already nails it — no App required. An App earns its keep only when the screen is genuinely interactive: multiple controls, live displays, several actions the user chooses between. A vending machine job does not need a control room.
An App's on-screen state is not a flow's results. An App's components
share a state while the user works — what row is selected, what data is
loaded, what a text box currently holds. That is a live, screen-level thing.
It is not the same as the results a flow's steps pass down its conveyor
belt (How Steps Pass Data). One is the App remembering what the user is
doing right now; the other is backend steps handing data to each other.
Confusing the two leads to looking for state in the wrong place.
Apps are for end users who need a screen. The reason to build an App is a real human who will sit in front of it and operate something. If your only user is a developer who runs the thing occasionally, the flow's run page is already enough — building an App for that audience is effort spent on a control room nobody will stand in.
A flow with its form is a vending machine — one job, one button; an App is a control room, build one only when one button truly isn't enough.
Your script already returns the answer. The last mile is letting somebody read it without logging into Windmill and clicking Run.
You have written the script. It runs, it is correct, and the answer comes back in the job result panel where you can see it perfectly well.
Now your director wants to look at it. Not to run it — to look at it, on a Tuesday, from a bookmark, without learning what a workspace is. That last stretch, from "the job produced the answer" to "a person read it," is the part nobody writes a ticket for, and it is where a lot of good work quietly stops.
Windmill can cover that stretch, and it does not need a web server to do it.
Think about a restaurant kitchen.
Most of the time, the way food reaches you is that a server walks into the kitchen, picks up the plate, and carries it out. That is the normal path: an authorized person goes to where the work happens, collects the result, and brings it to whoever wanted it. It works, and it requires that somebody knows their way around the kitchen.
Then somebody cuts a service window into the wall. Same kitchen, same cooks, same recipes. But now a plate can be set on the ledge and picked up from the street by a person who has never been inside and never will be. Nothing about the cooking changed. What changed is that the kitchen grew a way to hand things out.
An HTTP route is that window. Your script is the kitchen: unchanged, still doing exactly what it did. The route is the hatch cut into the wall, and the URL is the address on the street.
A route in Windmill is a trigger: a URL that, when someone requests it, causes something to happen. When you create one, you choose what sits behind the window, and this is the choice that matters most:
Intuition says the first one is for serving pages and the second is for APIs. That is the wrong split, and believing it will send you looking for a storage bucket you may not have. The static options are bound to object storage: their configuration requires a bucket key, so with no bucket there is no static route, no matter how simple your page is.
The runnable option has no such dependency, and it is not limited to returning JSON. On a synchronous endpoint — one where the caller waits for the answer rather than getting a job id — Windmill lets the script describe its own HTTP response through special keys in the returned object:
wm_content_type becomes the response's Content-Type headerresult becomes the response bodywm_status_code sets the status, wm_headers adds headersSo a script that returns {"wm_content_type": "text/html", "result": "<html>…"}
is not returning data about a page. It is returning the page. The browser gets
text/html, renders it, and the reader has no idea a Python function was
involved.
Follow one request. A browser asks for the route's URL. Windmill checks the
trigger, runs the script, and looks at the object that came back. It pulls
wm_content_type out to build the header and result out to build the body,
and sends both. The script never touched a socket and never imported a web
framework; it returned a dictionary and Windmill did the rest.
Notice where sync matters. On an asynchronous endpoint the caller gets a job
id immediately — there is no completed result yet, so there is nothing to shape
into a response, and the special keys have no meaning at all.
The whole mechanism, in one function:
import wmill
def main() -> dict:
html = wmill.get_variable("f/reports/latest_page")
return {
"wm_content_type": "text/html; charset=utf-8", # the header
"result": html, # the body
}
Point a route at it — method GET, target Runnable, request type sync —
and the URL renders a page.
Two things this example is quietly doing right, both worth copying.
It reads, it does not compute. A route runs on every page load. If the script measures something expensive, every visitor pays that cost and every visitor hammers whatever it measures. Put the expensive work in a scheduled job that writes the answer somewhere, and let the route do nothing but hand it over.
It sends a complete, self-contained page. A route returns one response. If your HTML links to a stylesheet, an image, or a script at another path, the browser will ask for those too, and nothing is listening there. Inline everything, or serve a page that needs nothing.
For binary content — a PDF, an image, a spreadsheet — sync endpoints serve the
result as text, so add wm_content_transfer_encoding: "base64" and return the
body base64-encoded. Windmill decodes it and serves the raw bytes with the
content type you gave.
1. "Serving a file" and "serving a page" are different doors. The route options named for static content require object storage; the option named for running code does not. If you have no bucket, the door you want is the runnable one — which reads backwards, because you are not really "running" anything, you are handing back a document. Pick the door by what it depends on, not by what it is called.
2. The special keys only work on a sync endpoint. Set the request type to
async and wm_content_type is ignored, because the response is a job id and the
result does not exist yet. The symptom is that your carefully built page arrives
as a small JSON object with an id in it, which looks like the script failed. It
did not; nobody waited for it.
3. A route bypasses the result viewer, and that is a feature. Windmill's job
result panel renders HTML but strips <style>, so a page previewed there comes
out unstyled and looks broken. That stripping belongs to the viewer, not to
Windmill, and a route does not go through the viewer. If your page looked plain
in the run panel, do not start rewriting the CSS — look at it through the route
first.
4. Public means public to whoever can reach the host. A route's authentication can be set to none, and on an internal server that means "anyone on the network," not "anyone on the internet." Both of those are much larger groups than "the person I meant to send this to." Decide who should see the content before you decide how convenient the link should be; convenience is easy to add later and impossible to take back.
5. The URL is part of the deliverable. Once somebody bookmarks it, the path belongs to them. Whether the workspace name appears in it is a toggle you set when creating the route, and changing your mind later changes the link. Decide it before you send the link, not after.
An HTTP route can run a script, and a script on a sync endpoint sets its own Content-Type — so the last mile to a real web page is one return statement, not a web server.
Your parameter has a default. The form says the field is required. Both of those are true at the same time, and the reason will change how you write every signature.
You wrote a perfectly ordinary signature. One parameter defaults to a string, a second to a constant defined at the top of the same file, a third to a small expression. Three defaults, all valid Python, all obviously correct.
Open the form and one field is pre-filled and the other two are marked Required — with a red asterisk, as though you had never written a default at all.
You did write them. Python knows about them. The form does not. This article is about the gap those two sentences describe, and about the second, quieter half of the problem: what actually reaches your function when somebody submits that form anyway.
Picture an order slip going to a warehouse.
Most lines are unambiguous. Quantity: 12. The picker reads the number, counts twelve, done. Nothing to interpret; the number is the instruction.
Now one line where, instead of a number, you wrote "the usual". To you that is not vague at all — it is precise, it has a value, and anyone in your office could tell you what it is. But the person picking the order has never worked with you, has no file to consult, and cannot leave the floor to find out. They cannot put "the usual" in a box.
Here is the part that matters. They do not guess. They do not go looking. They leave the line blank and the crate goes out with nothing in that compartment — not a smaller amount, not last month's amount. Nothing.
A literal default is a number on the slip. A constant, an expression, a function call — anything that has to be looked up to become a value — is "the usual". It means something to you and nothing at all to the reader.
To see why, you need one fact from Inputs Become Forms: Windmill builds
the form's schema from your main() signature.
The question is how it reads that signature, and the answer is: as text. Windmill parses your file, walks the syntax tree, and looks at each parameter's default. It never executes your module — and it must not. Running arbitrary user code just to draw a form would be slow and dangerous.
That constraint decides everything. A parser reading source can resolve a
literal, because a literal needs nothing outside itself: "prod" is the string
prod, 50 is the number 50, full stop. It cannot resolve a name, because
knowing what MY_CONST holds means executing the line that assigns it. It
cannot resolve a call or an expression, for the same reason.
So the parser does the only honest thing available: where it cannot read a value, it records that there is no default. And a parameter with no default is a required field.
That is the correction at the heart of this article, and it is worth stating flatly because the intuitive story is so appealing. The default is not evaluated once and frozen. It is not stale. It does not hold an old value. It is gone — never captured, never stored, absent from the schema entirely.
Three parameters, measured on a live worker, is the whole proof:
MY_CONST = "from_a_constant"
def main(
a: str = "a_literal", # a literal
b: str = MY_CONST, # a name
c: str = "x" + "y", # an expression
) -> dict:
return {"a": a, "b": b, "c": c}
The form that comes back:
a string [ a_literal ] <- pre-filled
b * string [ ] Required <- the default is simply not there
c * string [ ] Required
Now the second half, which is the part that bites. Submit that form with b and
c left blank and run it. The result:
{ "a": "a_literal", "b": "", "c": "" }
Read b again. Not from_a_constant. An empty string.
Python's default would have applied if main() had been called without that
argument. It was not. Windmill passes every parameter explicitly, using whatever
the form held, and the form held nothing. Your default did not lose a race
against the form; it was never in the race. The line you wrote to guarantee a
sensible value guarantees nothing at all.
The shape that fails, and it is the shape everyone writes:
PREFIX = "f/reports/page"
def main(prefix: str = PREFIX) -> dict: # field is REQUIRED; blank arrives as ""
return load(prefix) # load("") — not load(PREFIX)
The fix is two lines and it is mechanical. Put a literal in the signature so the schema can hold it, and resolve the real value in the body, where code actually runs:
PREFIX = "f/reports/page"
def main(prefix: str = "") -> dict: # a literal: optional, pre-fills as blank
prefix = prefix or PREFIX # the body decides what blank means
return load(prefix)
Now the parameter is optional, an empty submission is expected rather than surprising, and the fallback lives in the one place that is guaranteed to execute on every single run.
None works as well as "" and is often clearer about intent:
def main(prefix: str | None = None) -> dict:
prefix = prefix or PREFIX
Either way the principle is the same: the signature is read, the body is run. Anything that must be computed, looked up, or decided belongs on the side that runs.
1. "It has a default" and "the form has a default" are different facts. Your function is genuinely callable with no arguments — from a test, from another script, from a REPL. None of that reaches the form, because the form was built by reading, not by calling. The same file behaves one way under Python and another under the parser.
2. The default is not stale — it is missing. This distinction is the whole article. A stale value is wrong and present: you would see it in the field and could question it. A missing default produces a required field and, if submitted blank, an empty value. There is nothing to look at and nothing to suspect, which is why this costs people an afternoon.
3. A blank submission does not fall back to Python. Windmill calls main()
with every parameter supplied. An empty box becomes "", an argument that was
explicitly passed, so Python never reaches for the default. Any guard you wrote
assuming "if they leave it alone, my default applies" is inverted: leaving it
alone is precisely how you get nothing.
4. Empty is more dangerous than missing in a filter. An empty string does not
usually crash. It flows into a query, a path, a comparison. A condition written
as col = :p OR :p = '*' behaves very differently for "" than for a real
value, and a filter that quietly matches everything returns a confident, wrong,
much larger answer. See Shaping the Form for shaping inputs so blank is a
state you chose to handle.
5. Mutable defaults are the same rule, not a special case. x: list = [] is
a list display in the syntax tree, not a literal the parser can bank, so it is
dropped exactly like a name — required field, empty submission. Python's classic
shared-mutable-default trap is real and worth knowing, and it is a different
problem from this one. Default to None and build the container in the body,
which fixes both at once.
Windmill reads your signature as text, so only a literal default survives — anything else is dropped, the field turns required, and a blank submission arrives as empty, never as your default.
You declared three dependencies — your worker installed forty; the other thirty-seven are the ones that surprise you.
You opened your script, typed three lines into #requirements:, and saved.
Windmill resolved it, the worker built the environment, and your code ran. So
far so good. But if you go look at what the worker actually installed, you will
not find three packages. You will find forty. You declared three; the other
thirty-seven arrived on their own. And it is almost always one of those
thirty-seven — a package you have never typed, never imported, never heard of —
that breaks your run.
Think of throwing a small party and writing a guest list.
You wrote one name: pandas. One guest, one line. But pandas does not arrive alone. It cannot function without certain friends, so it brings them: numpy, python-dateutil, pytz, and a few more relatives. Those are its transitive dependencies — guests you did not invite, who came because someone you did invite cannot show up without them.
Usually this is fine. The extended family is well-behaved; they mingle, they help, you barely notice them. But invite a second guest with their own family, and the two families can clash in your living room. Two cousins want the thermostat at different temperatures and neither will budge — and you never invited either cousin. You only ever named two guests. The fight is happening between people whose names were never on your list.
There are two kinds of dependency, and the difference is the whole article.
A direct dependency is one you chose. You wrote import requests, or you
put requests in your #requirements: block. You know it is there because you
asked for it. A transitive dependency is one your dependencies chose. You
never imported it and never declared it; it is in your environment only because
something on your list cannot run without it.
When Windmill resolves your script (#requirements), it does not stop at
your three lines. It reads what those packages require, and what their
requirements require, all the way down — and the worker installs the entire
tree (The Worker). So your real environment is not your #requirements
list. It is the full resolved tree, and that tree is usually ten times bigger
than the list you typed.
The lock file is the honest guest list. Your #requirements block is what
you asked for; the lock is what you got — every direct package, every
transitive package, each pinned to one exact version. When you want to know
what is actually running on the worker, the lock is the document that tells the
truth.
The trouble starts when two branches of the tree reach for the same package
and disagree about its version. Picture two of your direct dependencies: one
needs numpy<2, the other needs numpy>=2. You never wrote numpy anywhere —
it is transitive to both. The resolver now has an impossible request: one
version of numpy cannot satisfy both constraints at once.
This is the moment the resolve fails, or quietly picks a compromise that leaves one of the two packages subtly unhappy. Either way, the conflict is on a name that never appears in your script.
Here is a #requirements: block. Three lines — the entire list you typed:
#requirements:
#pandas==2.2.2
#requests==2.32.3
#boto3==1.34.0
Three packages. But the worker does not install three. It installs the resolved set — the full tree those three pull in behind them:
# what actually lands on the worker (the locked set, abridged):
pandas==2.2.2 <- you declared this
numpy==1.26.4 <- transitive: pandas needs it
python-dateutil==2.9.0 <- transitive
pytz==2024.1 <- transitive
requests==2.32.3 <- you declared this
urllib3==2.2.1 <- transitive
certifi==2024.2.2 <- transitive
charset-normalizer==3.3.2 <- transitive
idna==3.7 <- transitive
boto3==1.34.0 <- you declared this
botocore==1.34.0 <- transitive
s3transfer==0.10.0 <- transitive
jmespath==1.0.1 <- transitive
# ...and more
You wrote three lines. The lock has a dozen-plus, and that dozen-plus is your
real runtime environment. Every name with <- transitive next to it is a
package you can break on, debug against, and version-conflict over — without it
ever appearing in the script you wrote.
Your environment is far bigger than your #requirements. What breaks at
runtime is often a dependency you never named — an ImportError or a crash
deep inside a package you have never opened. When you debug, do not stare
only at your three declared lines. Open the locked set and read the
whole tree; the culprit is usually a name you would not have thought to
look for.
Version conflicts come from transitive deps. Package A wants numpy<2,
package B wants numpy>=2. Neither numpy line is in your list, yet the
resolve either fails outright or picks a compromise version that leaves one
side subtly broken. The error message names numpy; your script does not.
That gap is exactly where the confusion lives.
Pinning only your direct deps does not pin the environment. You can pin all three of your declared packages to exact versions and still not have a reproducible environment — a transitive dep can float to a new version on the next resolve and change behavior under you. Real reproducibility is the whole resolved set locked, not just your three lines.
More direct deps means far more transitive deps. Every package you add drags its entire family in with it — and each family is another chance for two cousins to clash. A lean dependency list is not just tidiness or good taste; it is mathematically fewer packages, fewer versions, and fewer conflicts. The cheapest way to avoid a transitive-dep fight is to not invite the family that starts it.
Every dependency brings its family — your real environment is the resolved tree, not your #requirements list; debug and pin against the full lock.
Your code is correct — you proved it runs on your laptop, and it fails on the worker; it might not be the code.
Your code is correct. You did not just hope so — you proved it. You ran it on your laptop, watched it connect to the test database, watched the rows come back. It works. Then you deploy it, hit Run, and it fails.
So you go back to the code. You re-read the SQL. You check the connection string character by character. You add a retry. Nothing helps — because the code was never the problem. Before you debug it one more time, consider the possibility that has nothing to do with your code: it might not be the code.
Think of a phone call placed from two different buildings.
Your laptop and the worker do not sit in the same place — that is the whole lesson of The Worker, and it is worth holding onto here. They are in different buildings. From your desk, you pick up the phone and dial the test database's number. It rings, it picks up, you talk. Easy. You and that database are wired to the same internal switchboard — the office VPN — so the call goes straight through.
Now the worker, in its building, picks up an identical phone and dials the exact same number, using the exact same code to dial it. And it hears nothing. No ring, no pickup, just silence. The number is right. The dialing is right. But the worker's building has a different switchboard, and that switchboard has never been told it is allowed to connect to that number. Same number, same dialing, same everything you can see in the code — and a completely different result, because the call is leaving from a different building.
When a script reaches out to a database or an API, it is quietly doing two separate jobs at once. First, it runs code — it parses your query, builds the request, handles the response. Second, it opens a network connection — it asks the operating system to actually reach a host at some address and port. These are not the same job, and one being right tells you nothing about the other.
Code correctness and network reachability are independent. Your laptop, sitting on the office VPN, and the worker, deployed wherever it happens to be deployed, have completely different network vantage points — different firewalls in front of them, different routes out, different allowlists deciding who they may talk to. A database your laptop can reach in a heartbeat may be firewalled off from the worker entirely. Nothing about the host changed; nothing about your code changed. What changed is who is asking.
And here is the cruel part: when the network blocks the connection, the symptom
looks exactly like a code bug. You get a timeout, or a flat "connection
refused." It surfaces as an exception, with a stack trace, pointing at the line
where you called connect(). It feels like the code broke. But the code
never even got to run its logic — it was stopped at the door, before the first
line of your actual work.
"Works on my laptop" is a true statement. The trouble is that it proves the wrong thing. It proves your code is correct and that your laptop can reach the host. But your laptop is not where the job runs. The worker is the only environment that counts, and you have not tested it at all.
Picture the two places side by side: your laptop on the VPN with a clear line to the test database, and the worker somewhere else with a firewall standing between it and that same database. The code is identical in both. The only difference is the line each one has to cross.
Here is an ordinary database connection — the kind that lives in a thousand scripts:
import psycopg2
def main():
conn = psycopg2.connect(
host="test-db.internal.example.edu",
port=5432,
dbname="nsc_test",
user="nsc_reader",
password="...",
)
# ... your actual query logic ...
Read it closely and look for the bug. There isn't one. This exact code, byte for byte, runs perfectly on your laptop and fails on the worker. The code is not what differs between the two runs — where it runs from is. On your laptop the call leaves the VPN and arrives. On the worker the call leaves a different network and is blocked before it arrives.
So the fix is not in this file. No edit to these five lines will help, because
these five lines are already correct. The fix is a network change — a
firewall allowlist rule that permits the worker's address to reach
test-db.internal.example.edu:5432 — or routing the job through a host the
worker can already reach. Either way, the change happens outside Python.
Your debugging instincts were trained on a world where your machine and the running machine were the same machine. The worker breaks that assumption, and four habits go wrong with it:
"It works on my laptop" proves your laptop, not the worker. That sentence is true and useless at the same time. It confirms two things — your code is correct, and your laptop's network can reach the host — and it says nothing about the worker, which is the only environment that actually counts (The Worker). A green run on your machine is not evidence about a red run on the worker.
A timeout or "connection refused" is a network answer, not a code answer. Read the error before you act on it. A timeout means the worker reached out and got silence; "connection refused" means it reached the host but the door was shut. Both of them say could not reach — not logic wrong. When you see those words, do not open your SQL again. The error is already telling you it never got that far.
You cannot fix a firewall with Python. This is the hard one to accept. No retry loop, no clever library, no timeout tuning, no amount of code will talk a blocked host into answering. A wall is a wall. The fix lives outside your script entirely — a network change, or routing the job through a host the worker can reach. Hours spent rewriting code against a firewall are hours spent against a thing code cannot touch.
Test from the worker, not from your laptop. The only honest test is one that runs where the real job runs. Write a tiny "can I connect?" smoke test — open the socket, report success or failure — and run it as a Windmill job so a worker executes it. That tells you the truth. The same test on your laptop tells you a comforting lie. If the job is scheduled, run the smoke test on the same schedule path so it leaves from the same place (From crontab to Schedules).
Code correctness and network reachability are two different things — 'works on my laptop' only proves the first; when a connection times out, suspect the network.
You built it, you saved it, you can see it in the list — and the URL says it does not exist.
You filled in the form carefully. Path, method, the script it should run, the authentication. You pressed Save and it saved — no red text, no complaint. You can open the list and see it sitting there with the name you gave it.
Then you open the URL and Windmill tells you:
Not found: Trigger not found at name /bcm
Not found. The thing you are looking at, in a list, on your screen. This article is about that sentence, because the sentence is lying to you — not on purpose, but in a way that will cost you an hour if you believe it.
Picture a shop on the evening before it opens.
Inside, everything is finished. The shelves are stocked, the register works, the floor is swept, the staff know their jobs. If you were standing inside you would say, without hesitation, that this shop is ready. It exists. It is complete.
From the street, none of that is visible. The sign above the door is dark. A person walking past does not think "this shop is closed" — that would require noticing the shop at all. They think nothing. As far as the street is concerned, there is no shop here.
Being built and being open are two different states, and only one of them is visible from outside. A trigger is exactly this. Creating it builds the shop. Enabling it turns on the sign. Windmill does the first when you press Save. It does not do the second.
A trigger is two things that intuition fuses into one.
The first is a definition stored in your workspace: this path, this method, this runnable, these settings. That is what Save creates, and it is what the list shows you. It is a row in a table.
The second is a live registration in the router — the part of Windmill that receives an incoming request and decides what to do with it. The router does not read the whole table of definitions. It works from the set of triggers that are currently switched on.
A newly created trigger is in the first state and not the second. It is a complete, valid, correct definition that the router has never been told about. So when a request arrives, the router looks through what it is listening for, finds nothing matching that path, and answers honestly from its own point of view: there is no trigger here.
The message is accurate about the router and misleading about your workspace. Both things are true at once, which is exactly why the error is so effective at sending you in the wrong direction.
Three quite different situations produce that same sentence, and telling them apart is the whole skill:
One message, three causes. The list tells you whether you are in case 1. The toggle tells you whether you are in case 2. The trigger's own page, which displays the URL it actually answers on, tells you whether you are in case 3.
There is no code in this one, and that is the point — nothing you can write will fix it. The fix is three checks, in this order, each one ruling out a whole class of cause:
1. Open the triggers list.
Not there? -> it never saved. Look for a required field
hiding in the collapsed "Advanced" section.
There? -> go to 2.
2. Look at the enable toggle on the trigger.
Off? -> that is your answer. Turn it on.
On? -> go to 3.
3. Read the URL shown on the trigger's own page.
Different from the one you typed?
-> use the one it shows. Never hand-assemble it;
the workspace prefix is a setting, not a rule.
Do them in that order and the answer arrives in under a minute. Do them in any other order — or skip to reading your script — and you can lose an afternoon to a switch.
1. Save is not Enable, and only one of them is loud. Every other part of the form validates: leave a required field blank and it tells you. The enable state is not a validation problem, so nothing warns you. The form's job ended when the definition was valid; nobody's job was to ask whether you wanted it listening.
2. The error describes the router's world, not yours. "Not found" is what the router can truthfully say, because the router genuinely does not have it. Most error messages describe the system's state; you read them as descriptions of your state. When a message contradicts something you can see with your own eyes, the useful question is not "am I wrong?" but "which component is speaking, and what does it know?"
3. A disabled trigger is indistinguishable from a wrong URL, from outside. Both give the same sentence. So does a typo in the path. The list and the toggle are the only two places that tell them apart, and neither is where the error points you. When the same symptom has three causes, stop reasoning and start eliminating.
4. This applies to every trigger type, not just routes. Schedules, webhooks, and the rest all share the shape: a definition that exists and a registration that may or may not be live. A schedule that never fires is the same bug wearing a different hat, and it is worse there, because a schedule that does not fire produces no error at all — just silence, which nobody investigates until the report they expected never arrives.
5. The default is off for a good reason, and it is still a trap. You do not want a half-configured public URL going live the moment you press Save; off-by- default is genuinely the safer choice. Both things can be true: the default is right, and the error message it produces is wrong about why. Good defaults do not excuse misleading errors.
A new trigger is created switched off, and the error it gives you says "not found" instead of "not on" — so check the toggle before you doubt anything else.
A variable will hold your API key, your file path and your feature flag. Hand it a document and you find the wall.
Variables are the most obliging thing in Windmill. A connection string, a folder path, a threshold, a flag somebody flips twice a year — write it, read it, done. Nothing about them suggests a size.
So the first time you try to keep something substantial in one, you are not expecting resistance. And what you meet is not a helpful error explaining the design. It is a character counter, quietly ticking toward a number you did not know existed.
An envelope and a filing cabinet are both places you put paper, and they are not interchangeable.
An envelope is for something small with a name on the outside. A note, a key, a card. Its whole virtue is that it is trivial to label, trivial to hand to someone, trivial to replace. Nobody designs an envelope wondering how many pages it should hold, because the answer is "a few, obviously."
A filing cabinet is for documents. It has structure, an index, and the expectation that things inside it are large.
If all you have is envelopes and you need to move a hundred-page report, you do not give up. You do what people have always done: you compress the thing as far as it will go, split what is left across several envelopes, number them, and put an index card in front saying how many there are. That last part is not tidiness. Without it, whoever opens the drawer has no way to know they are holding all of it.
A Windmill variable is an envelope. This article is the numbering system.
A variable holds a string, and the store enforces a size limit on that string.
The exact ceiling belongs to your instance and shows up in the editor as a live
counter — on the instance this article was measured against, it reads
0/10000 characters.
Ten thousand characters is generous for what variables are for and small for a document. Measured on a real page: 192,100 characters of HTML. That is twenty envelopes, and none of them would hold a complete tag, which means every cut falls in the middle of markup.
Two moves fix it, in this order.
Compress first. Text compresses hard. That same 192,100-character page came down to 38,528 characters after gzip and base64 — from twenty envelopes to five. Base64 is not decoration: the result of compression is arbitrary bytes, and a variable holds text, so the bytes need a text-safe spelling to survive the trip.
Then split, and index. Store the compressed text as numbered pieces plus one small manifest — a piece that is not content, only a description of the content: how many pieces there are, and a checksum of the original. The reader opens the manifest first, learns how many to expect, collects exactly that many, joins them, and verifies.
The manifest is what turns a pile into a document. Without it, a reader has to guess where the pieces stop, and a reader that guesses will eventually guess wrong at the worst possible moment.
Now the timing, which is the part that separates a design that works from one that mostly works.
Writing is not instantaneous. If your job replaces five pieces one at a time, there is a window in which some pieces are new and some are old. Anyone reading during that window gets a mixture, and a mixture of two different compressed streams is not a slightly wrong document — it is not a document at all.
So write the pieces first and the manifest last. The reader keys off the manifest, which means that until that final write lands, the manifest still describes the previous document. The window shrinks to a single write, and the checksum turns whatever is left into a clear, named error instead of a wrong page.
Writing:
import base64, gzip, hashlib, json, wmill
PREFIX, CHUNK = "f/reports/page", 8000
def publish(html: bytes):
packed = base64.b64encode(gzip.compress(html, 9, mtime=0)).decode("ascii")
parts = [packed[i:i+CHUNK] for i in range(0, len(packed), CHUNK)]
# Prove it survives the round trip BEFORE replacing anything good.
if gzip.decompress(base64.b64decode("".join(parts))) != html:
raise RuntimeError("did not round-trip; nothing was written")
for i, part in enumerate(parts): # pieces first
wmill.set_variable(f"{PREFIX}_{i:03d}", part)
wmill.set_variable(f"{PREFIX}_manifest", json.dumps({ # manifest last
"chunks": len(parts), "sha256": hashlib.sha256(html).hexdigest()}))
Reading:
def load() -> bytes:
m = json.loads(wmill.get_variable(f"{PREFIX}_manifest"))
parts = [wmill.get_variable(f"{PREFIX}_{i:03d}") for i in range(m["chunks"])]
blob = gzip.decompress(base64.b64decode("".join(parts)))
if hashlib.sha256(blob).hexdigest() != m["sha256"]:
raise RuntimeError("does not match the manifest — a write is in flight")
return blob
Note mtime=0 in the compress call. Without it gzip stamps the current clock
into its own header, so compressing the same document twice produces different
bytes. That single argument is what makes a re-run idempotent, and it is the
subject of the third gotcha below.
1. Too big fails at write time; wrongly assembled fails at read time, silently. Exceeding the limit is the friendly failure — something stops you. The dangerous one is a document that reassembles into garbage: gzip may or may not notice, and a browser handed a truncated page renders a blank screen with no error anywhere. Verify with a checksum; do not rely on the decompressor to be suspicious on your behalf.
2. Order is not obvious to the reader, and it must be. Numbered names
(_000, _001) exist so that reassembly is arithmetic rather than inference.
If you name pieces after their content, or sort them as text without padding,
_10 sorts before _2 and you have invented a bug that only appears on the
eleventh piece.
3. Compression is not deterministic unless you make it so. gzip writes the
current time into its header. Compress the same input twice and you get two
different outputs. This is harmless right up until someone re-runs the writer
while another person is halfway through a manual paste — now some pieces are
from generation A and some from B, and they will never fit together. mtime=0
removes the problem at the source.
4. A larger ceiling would not change the design. The instinct on meeting the limit is to look for a way around it. But splitting is not a workaround for a number being too small; it is the correct shape for content that grows. A page that fits today will not fit forever, and a design that depends on fitting is a design with an expiry date nobody wrote down.
5. Secret variables are a bad envelope for this. Marking these secret buys nothing — compressed HTML is not a credential — and costs two real things: you can no longer read the value back to check what you stored, and every read generates an audit-log entry. Multiply that by five pieces on every page load and you have built a noise machine. Save secrecy for things that are actually secret; see Variables vs Secrets.
A variable is a labelled envelope, not a filing cabinet — past about 10,000 characters you must compress, split and record how many pieces there are, or you will serve half a document with no error at all.
It ran perfectly for eight months. Then the person who wrote it changed roles, and nothing worked — with no error anyone could find.
The schedule had run every Monday at five in the morning for eight months. The report landed, people read it, nobody thought about it — which is the highest compliment a scheduled job can receive.
Then the person who built it moved to another team. Nothing was deleted. No permission was revoked. No code changed. And the report stopped being right, in a way that took three weeks to notice and another two to explain.
The cause was in the first two characters of a path.
Think about where things live in an office.
There is the drawer in your desk. It has your name on it, and everything about that is correct: your things, your drawer, easy to reach, nobody else rummaging. When you move desks, the drawer's contents move with you, because they were always yours. Nobody would call that a bug. That is what a personal drawer is for.
Then there is the shared cabinet in the corridor, labelled by department. Anyone on the team opens it. When someone leaves, nothing in it moves, because nothing in it was theirs. The cabinet belongs to the work, not to a person.
Both are storage. Both hold paper. Choosing between them is not about convenience — it is a statement about who the thing belongs to, and that statement is enforced later, at the worst possible moment, by reality.
In Windmill this choice is not a setting or a checkbox. It is the first segment of the path, and you make it every time you name something, whether or not you realise you are making it.
Every path in Windmill begins by declaring a kind of owner:
u/alice/… — the user namespace. This belongs to the person named
Alice. Permissions follow Alice.f/payroll/… — the folder namespace. This belongs to the folder
payroll, which has its own members and its own permissions. People come and
go from that folder without anything inside it moving.Intuition reads u/alice/daily_report as a path that happens to start with
someone's name, the way /home/alice/ does on a laptop — a location that is
merely conventionally hers. That is the wrong model. It is not a location that
belongs to Alice by convention; it is a location that is Alice's, structurally,
and Windmill's permission system knows it.
This applies to everything that has a path: scripts, flows, variables, resources, triggers, schedules. And it is why two things that look like naming preferences are actually architecture decisions:
u/alice/api_key is readable by jobs running as Alice. A job
that later runs as someone else may simply not find it.None of that is a misconfiguration. It is the system doing precisely what the path told it to do.
Follow the failure, because it is quieter than you expect.
A job writes its output to a path under the user who set it up. A second job reads from that path. While the same person owns both, everything lines up and the pair works perfectly — for months, which is the problem, because months of success is how something stops being questioned.
Then ownership changes: the schedule is re-saved by someone else, or the original account is deprovisioned. The writer now resolves to a different place. Nothing errors. The writer writes successfully — to the new location. The reader reads successfully — from the old one, where the last good copy still sits.
The pipeline reports success at every step and quietly serves a frozen answer. That is the worst failure shape there is: no exception, no red job, no alert, and a number on a page that people keep trusting.
The change is small enough to look like a style preference, which is exactly why it gets skipped:
# Tied to a person. Works until the person changes.
VAR = "u/alice/report_page"
# Tied to the work. Survives any handover.
VAR = "f/reporting/report_page"
The same rule governs object-storage keys, where it is enforced by rules that mirror these namespaces. Permission rules commonly read like this:
u/{username}/**/* -> allowed
f/{folder_write}/**/* -> allowed
**/* -> deny
Read the first line carefully. {username} is resolved at run time, to
whoever is running the job. It does not mean "the user who wrote this code"; it
means "whoever is here right now." Write to u/{username}/output.csv from a
schedule and the file lands wherever the current runner's name points, which is
a different place after a handover — and the deny-all at the bottom means the
next attempt might not fail loudly either, it might just be somewhere else.
The practical test, before you name anything:
If the person who is setting this up left tomorrow, should this keep working?
Yes → f/. No → u/ is genuinely the right answer, and personal scratch work
belongs there.
1. u/ looks like a home directory and behaves like an owner. Every
developer has years of muscle memory saying that a path starting with a username
is just a place. Here it is a claim about who the thing belongs to, and the
permission system acts on that claim. The syntax borrows from filesystems; the
semantics do not.
2. The failure arrives months later, attached to a personnel change. Nobody debugging a stale report on a Tuesday is thinking about a role change from six weeks ago. The cause and the symptom are separated by so much time and so many unrelated events that the connection is nearly invisible. Fix it while naming things, because you will not find it while debugging them.
3. It succeeds loudly and is wrong quietly. Neither job fails. Both logs are green. The output exists, is readable, is well-formed, and is old. Green logs are evidence that each step did what it was told — never that the steps were told the right thing.
4. Permissioned-as is the same decision wearing different clothes. You can
put every path under f/ and still leave a schedule or route permissioned as a
person. The storage is now shared and the execution identity is not, so the
whole chain still hangs off one account. Check both. A URL that leadership
bookmarks should not depend on one employee's account existing.
5. Moving it later is real work, so decide it at creation. Changing a path
means recreating the object, migrating whatever is stored under it, and updating
every reference. It is not hard, but it is a task with a risk of loss, and it
always arrives at a worse moment than the ten seconds it would have cost to type
f/ the first time.
Every path in Windmill starts with who owns it: u/ ties the thing to a person
and f/ ties it to a folder, so anything that must outlive one employee belongs
under f/.
You pushed three files. The workspace had two hundred. Only three survived.
Every developer alive has an instinct about the word push. You push your commits. You push an image. You push a file to a server. In every one of those, push means add: the destination ends up with what it had, plus what you sent.
wmill sync push does not mean add. It means make the remote look like this.
And the moment you understand that sentence, a second one follows immediately
and much less comfortably: everything the remote has that your folder does not
is, by definition, something that should no longer be there.
Two pieces of paper travel with deliveries, and confusing them is expensive.
A packing list says what is in the box. Take the items out, put them on the shelf, done. Nothing on that shelf is at risk from a packing list. It only ever adds. That is what people picture when they hear "push."
A stock reconciliation sheet is a different document with a similar look. It describes what the shelf should contain, in full. The clerk walks the aisle with it and makes reality match: missing items get added, and anything on the shelf that is not on the sheet gets pulled. The sheet does not have to mention an item to remove it. It removes it precisely by not mentioning it.
Held at arm's length the two sheets are indistinguishable. The difference is not in what they say. It is in what silence means. On a packing list, silence means "not in this box." On a reconciliation sheet, silence means "should not exist."
wmill sync push is a reconciliation sheet.
Sync's model is that a folder on your machine is the intended state of some region of the workspace. Push compares the two and issues whatever operations close the gap: creates, updates, and deletes.
That is not a flaw. For its intended use — one repository holding one workspace, versioned in git, deployed by CI — it is exactly right. When someone deletes a script in the repo and merges, they mean it should go away, and a tool that only ever added would leave the workspace accumulating dead objects that no one dares remove by hand.
The trouble is that the same tool, pointed at a folder that is only part of a shared workspace, reads that partial folder as the whole intended truth. Your colleagues' work is not something push is choosing to leave alone. It is something push has been told does not belong.
Which means the single most important line in your sync configuration is not the credentials or the workspace name. It is the one that decides how much of the workspace this folder claims to describe:
includes:
- f/my_project/** # this folder speaks for exactly this path
Read that as a scope of authority, not a convenience filter. Everything inside it, this folder now owns and may delete. Everything outside it is invisible and therefore safe.
Picture a workspace with two projects in it, yours and someone else's ten-step production pipeline. Your laptop has a folder containing only yours.
With includes pinned to your path, push compares your folder to your region.
The other project sits outside the comparison entirely. Nothing about it is even
considered.
Widen includes to the whole workspace and the comparison changes shape. Now
push is comparing your small folder against everything, and everything it does
not find locally is a difference to be closed. The other project is not
attacked; it is simply absent from the intended state, and push closes that
gap the only way it can.
Same command, same folder, same intent. One line decides whether the other team still has a pipeline on Monday.
The habit, in full. It is three lines and it is not optional:
wmill sync push --dry-run # 1. ALWAYS. Then actually read it.
wmill sync push # 2. only after 1 said what you expected
And the configuration that makes step 1 boring instead of frightening:
includes:
- f/my_project/** # the blast radius, stated on purpose
skipVariables: true # things whose live value is the truth
skipResources: true # never let a local copy overwrite them
skipSecrets: true
Those skip lines deserve a moment. A script is code and belongs in git. A
variable is state — the current value is the truth, and it lives on the server.
Sync a variable and a stale local copy can overwrite a value someone changed
yesterday. Worse, if what lives there is something you cannot regenerate, a
history, a counter, an accumulated snapshot, then a push can destroy the one
thing in the project that no amount of re-running will bring back.
How to read a dry run, in one rule:
If any line names a path you did not intend to own, stop.
Not "does the diff look roughly right." Read the paths. A destructive plan and a harmless one look almost identical at a glance, because most of both is additions.
1. Push means add everywhere else in your career. git push, docker push, scp. Every one of them is additive at the destination. This one is a mirror operation that happens to share the verb. Nothing in the command name warns you, and your hands will type it faster than your head can reconsider.
2. Silence is an instruction. The dangerous content of your local folder is not what is in it. It is what is missing from it, because absence is how you say "delete." This is why a partial checkout is more dangerous than an empty one, and why "I only changed one file" gives no protection at all.
3. A shared workspace changes the meaning of a wide scope. If the workspace
is yours alone, a broad includes is merely honest. Put one more team in that
workspace and the same line becomes a claim over their work. The config did not
change; the world around it did. Re-read includes whenever you learn that
somebody else is in there.
4. The dry run only helps if you read the paths. It is easy to run
--dry-run, see a wall of output, register "looks like my stuff," and proceed.
The failure mode is not skipping the dry run. It is skimming it. Look for paths
you do not recognise, and treat a single one as a full stop.
5. Code is replaceable; state is not. A deleted script is recoverable from git in seconds. An overwritten variable that held the only copy of something — a history, a counter, a measurement of a moment that has passed — is gone in a way no repository can fix. That asymmetry is the entire reason to skip variables and secrets, even when syncing them would be convenient.
sync push makes the remote match your folder, so anything the remote has and
your folder lacks is a deletion — pin includes to your own path and read the
dry run every single time.
The one line whose job was "this runs anywhere" is the only line that runs in exactly one place.
There is a line many careful Python programmers write without thinking, because it is the responsible thing to do:
HERE = os.path.dirname(os.path.abspath(__file__))
It means: do not assume the working directory, figure out where I actually live. It is the line you add so the script works from anywhere. It is portability, written down.
On a Windmill worker it can raise NameError and take the whole job with it —
which makes it, with some irony, the least portable line in the file.
Think about two ways the same letter can reach you.
One way, it arrives in an envelope. You hold the paper, and you also hold the postmark, the return address, the stamp. The message and its origin came together as one physical object. If you want to know where it came from, you turn it over. Of course you can.
The other way, someone reads it to you over the phone and you write it down. Every word is identical. The message is complete and correct. But there is no envelope, because there was never a physical object — the words were handed to you directly. Asking "what does the postmark say" is not a hard question here. It is a question with no referent. There is no envelope to turn over.
__file__ is the postmark. It exists when Python opened a file to get your code.
It does not exist when your code was handed to Python some other way — and
"handed over some other way" is a completely normal thing for a platform to do.
__file__ is not a Python built-in in the sense of len or print. It is a
module attribute that the loader sets, and it is set by loaders that loaded
your code from a path — because for those, a path is the honest answer.
Windmill's job is to take a script that lives in a workspace, resolve its
relative imports against workspace paths, and execute it on a worker. Under that
arrangement a module can be loaded without a filesystem path being the truth of
where it came from. So the loader has nothing meaningful to put in __file__,
and does the correct thing: it does not invent one.
Now the part that turns an inconvenience into an outage. That line usually sits
at module level, near the imports, because that is where constants go. Module
level code runs at import time, before any function is called. So the
NameError fires while the module is still being loaded — before main() runs,
before your logging is set up, before your carefully written preflight check has
a chance to report anything at all.
The script does not fail doing its work. It fails on the way in, and the log
shows a bare NameError: name '__file__' is not defined with no context, from a
line whose entire purpose was to make the code robust.
Compare the two lifetimes.
On a laptop: Python opens report.py, sets __file__ to its path, executes the
module top to bottom, and calls main(). Your line resolves. Everything is
fine, and stays fine, forever — which is exactly why nobody suspects it.
On the worker: the loader resolves the module through workspace paths, executes
the module top to bottom, and reaches your line. There is no __file__. The
exception propagates out of the import, the job dies, and none of the code you
wrote to explain failures ever runs, because it had not been defined yet.
This is a general shape worth naming: a module-level statement that assumes the environment is a promise you make at import time and cannot take back. Every assumption you push up to module scope is one you have decided to bet the whole job on, before you have any ability to report what went wrong.
The trap:
import os, sys
HERE = os.path.dirname(os.path.abspath(__file__)) # NameError on import
sys.path.insert(0, os.path.dirname(HERE))
The fix is to ask instead of assume, and to have a real answer for "not available":
import os, sys
_F = globals().get("__file__") # ask; never assume
HERE = os.path.dirname(os.path.abspath(_F)) if _F else os.getcwd()
# Only extend sys.path when there is a real file to be relative TO.
# Skipping is honest; inventing a directory would not be.
if _F:
for _p in (HERE, os.path.dirname(HERE)):
if _p not in sys.path:
sys.path.insert(0, _p)
globals().get("__file__") is the whole trick: it turns "a name that may not
exist" into "a dictionary lookup that may return None", which is an ordinary
condition you can branch on rather than an exception you have to survive.
And notice what the fallback does not do. It does not fabricate a plausible
directory. When there is no file, the sys.path manipulation is skipped
entirely, because on the worker there is nothing to be relative to — the imports
resolve by workspace path instead. A fallback that invents a path would keep the
script running while quietly making it wrong, which is worse than the crash.
1. It is not a Python guarantee; it is a loader courtesy. __file__ feels
as fundamental as __name__, and in a decade of scripts it has never once been
missing. That track record is about how you have always run Python, not about
the language. Anywhere code is executed from a string, a database, a notebook
cell, or a platform loader, the postmark is optional.
2. The most portable-looking line is the one that assumes the most. The
whole reason people write abspath(__file__) is to stop assuming things — to
stop depending on the working directory. It replaces a small assumption with a
larger, better-hidden one: that a file exists at all. Robustness that is aimed
at one environment can be fragility aimed at every other.
3. Failing at import time silences your diagnostics. This is the part that costs the hours. Teams build careful preflight checks that name the failing layer, and then a module-level line dies before any of it loads. If a check matters, nothing above it in the file may assume anything it is meant to check.
4. It works forever on the laptop, which is why it ships. There is no test that fails, no review comment, no warning. The behaviour is correct in every environment the author can see. The bug is created by the difference between environments, so it can only be found by running in both — which is the argument for one file that genuinely runs in both places.
5. The cousin of this bug is the relative path. open("data/config.json")
resolves against the working directory, which on a worker is not where you
think. Same family, same cause: a filesystem assumption that happened to hold
where it was written. If you need a file to travel with your code, it is not a
file — it is a variable, a resource, or a shared module.
A script on Windmill is loaded as a module, not opened as a file, so __file__
may not exist at all — ask for it with globals().get("__file__") and have a
real answer for None.
The file is listed, with a usage count, on a page called Assets. The write that was supposed to create it failed twenty minutes ago.
You are debugging a storage problem. The write threw an error you did not understand, so you go looking for evidence about whether anything is working at all.
You open the Assets page. There is your file, listed by its exact path, with a small note saying 1 usage. Relief. Storage is clearly configured, the file clearly exists, so the error must be something subtler — a permission, a timing, a race.
You have just been misled by a screen that never claimed to know anything about storage. The next three hours are the price of that misunderstanding, and the mistake is not carelessness. It is a reasonable inference from a page whose name promises one thing and whose contents are another.
A library has a card catalogue.
Pull a drawer and find a card: title, author, shelf number, typed neatly. That card is a real, accurate, useful record. It tells you the library knows about this book, where it would be, and who asked for it to be listed.
It does not tell you the book is on the shelf. Someone may have borrowed it, misfiled it, or ordered it and had the order cancelled. The card was created by somebody writing down an intention. Whether the object arrived is a completely separate fact, discoverable only by walking to the shelf and looking.
Nobody confuses these two in a library, because the drawer and the shelf are in different parts of the room and look nothing alike. In a web interface they are two tabs in the same sidebar, styled identically, and you cannot see which one walked to the shelf.
Panels in a platform are built from different sources, and the source is almost never shown.
Some panels are built by reading your code. The platform parses your scripts and flows, finds the paths they reference, and lists them. That is a static analysis of intent. It is genuinely useful: it answers "what does my code touch, and from where," which is a real question. It runs whether or not the storage behind those paths exists, because it never asked. A path that appears in a line of code that has never successfully executed appears here exactly the same as one written a thousand times.
Other panels are built by querying the store. Those list what is actually there, because producing the list required a round trip to the thing itself.
The two look the same and mean opposite things when they disagree. And "1 usage" is the tell, if you notice it: usage is a property of your code, not of a stored object. A bucket does not know how many scripts mention a key. Only a code index can count that, and a panel that can count it is a panel that read code.
There is a second layer of the same confusion, and it is worth naming because it
bites in the same hour. Platforms accumulate features with adjacent names —
storage, volumes, assets, artifacts. A reading like 0 volumes on one page says
nothing about object storage, because Volumes is a different feature with its
own settings. Meanwhile the honest signal for whether a bucket has ever received
a byte is often unglamorous and elsewhere: a size that reads 0.0B used.
The habit worth building is a single question, asked of any screen you are about to treat as evidence:
What did this panel have to read in order to draw itself?
If the answer is "my code," it can tell you about intent, references and usage, and it cannot tell you whether anything exists.
If the answer is "the store," it can tell you what exists, and it will be empty or slow or error when the store is unreachable — which is itself information.
When two panels disagree, the tiebreak is not which one is more prominent or more recently updated. It is which one had to touch the thing. A panel that never made a network call cannot testify about a network.
The way out of this class of confusion is not to squint harder at a screen. It is to ask the store a question whose answer only the store can produce, and to do it from a script rather than a UI:
import wmill, datetime
def main() -> dict:
"""Prove the store is reachable and writable. Not a UI reading — a round trip."""
key = "f/my_project/_probe.txt"
stamp = datetime.datetime.now().isoformat() # unique, so a stale
wmill.write_s3_file(key, stamp.encode()) # read cannot fool us
back = wmill.load_s3_file(key).decode()
return {"wrote": stamp, "read_back": back, "ok": back == stamp}
That is fourteen lines and it settles the question permanently. It writes a value
that cannot already exist, reads it back, and compares. If it returns ok: true,
storage works and you can stop wondering. If it raises, you have a real error
message naming a real layer instead of a screenshot to interpret.
Note the stamp. Writing "test" and reading "test" proves nothing if a file
called _probe.txt was already there from last year. A probe whose success can
be faked by history is not a probe.
1. The name of a panel is marketing, not a schema. "Assets" sounds like an inventory of things you own. It is an index of things you mentioned. Nobody lied; naming is hard and the word is defensible. But you cannot infer a data source from a label, and the label is all the UI gives you.
2. Presence is not proof of existence, and "1 usage" is the giveaway. A count of usages is a fact about code. If a panel can tell you a number like that, it read your scripts. Learning to spot that one tell converts this whole class of confusion into a two-second read.
3. Adjacent features borrow each other's words. Volumes, storage, assets, artifacts. A zero on one page is not a zero on another, and the reading you happen to find first is rarely the authoritative one. When a number surprises you, find the settings page that owns that feature before drawing a conclusion.
4. This is where a wrong conclusion gets written down. The real cost is not the confused hour. It is that somebody records "storage is configured" in a ticket or a runbook, and the next person inherits it as established fact and never re-checks. A finding that came from a screenshot should be labelled as such, or replaced by a probe before it is written down.
5. A screenshot is not a measurement. This is the general form, and it is worth carrying beyond this platform. A UI is a rendering of some query somebody else wrote, against a source you cannot see, at a time you did not choose. When the answer matters, ask the system a question yourself and keep the code that asked it — so that six months later the evidence can be re-run instead of re-interpreted.
Some panels list what your code mentions, not what the system holds — before trusting a screen as evidence, ask what it read to build itself.
The deploy went green. Every check passed. The job died three minutes later asking for a package nobody packed.
The deploy is green. Not "probably fine" green — green. The dependency step ran, found nothing to complain about, and finished. Whatever safety net your platform offers, it was extended and it held.
Three minutes into the first real run, the job dies:
ModuleNotFoundError: No module named 'oracledb'
Nothing is broken. Nothing is misconfigured. The deploy was right about everything it looked at, and it did not look at the thing that mattered. This article is about the gap between those two sentences.
Think about walking through customs.
You fill in a declaration card. The officer reads the card. If the card says something that needs attention, you get attention. If the card is clean, you walk through. The officer, in the ordinary case, does not open your suitcase. That is not laziness — checking every bag by hand does not scale, and the declaration is the mechanism the whole system is built on.
So the checkpoint is not measuring what you are carrying. It is measuring what you wrote down. Those two are usually identical, which is exactly why the system works and why nobody thinks about the difference.
Now imagine you carry something perfectly legal and simply do not mention it, because it did not occur to you that it counted. Customs is uneventful. The card was clean, the officer was satisfied, and you walked through with something the system never accounted for. The problem, if there is one, surfaces much later and somewhere else, with nothing about it pointing back to a border you crossed without incident.
Windmill's dependency step is that officer. It reads your declaration. It does not open the bag.
Two different things read your script, at two different times, in two completely different ways.
At deploy time, the dependency resolver reads your file as text. It parses the source and looks for import statements to build the list of packages the worker will need. It is a static read: it never executes a line, so it can only know what is written plainly where it looks.
At run time, the Python interpreter reads your file as a program. It executes it, and an import inside a function body happens the moment that function runs — which may be minutes later, or only on Tuesdays, or only when a particular branch is taken.
Intuition merges these into one idea called "the code," and then assumes that anything the interpreter will eventually need is something the resolver already knew about. It is not. The resolver saw the file the way a proofreader sees it. The interpreter sees it the way a performer sees a score.
So an import written at the top of the file is a declaration. An import written inside a function is an instruction that will be carried out later — invisible to a reader that is not running anything.
Here is what makes this one genuinely hard rather than merely obscure: the pattern that triggers it is often deliberate and correct. Imports get moved inside functions for real reasons — so one file can run on a laptop that does not have every package, so start-up stays fast, so an optional feature does not force a dependency on everyone. Those are good instincts. This is not a careless mistake meeting a strict tool; it is two reasonable ideas that are individually right and jointly broken.
Put the two timelines side by side and the shape of the trap is obvious.
The deploy timeline is short and ends in a verdict. Parse the source, collect the imports it can see, resolve them, report. Every step succeeds, because nothing in it is wrong. A resolver that found no third-party imports has not failed — it has correctly reported what a static read of your file contains.
The run timeline starts later and reaches further. It executes the module,
calls main(), enters a branch, and there hits an import for a package that was
never in the list, because it was never in the declaration.
The distance between those two moments is the whole problem. A failure that happens right after a check gets connected to the check. A failure that happens minutes later, in a different system, under a different log, does not — and the green deploy sits in everyone's memory as evidence that this part was verified.
The shape that fails, and it looks careful:
def main(term: str) -> dict:
import oracledb # kept local so this file also runs on a laptop
import openpyxl # ...where neither package is installed
...
Nothing here is sloppy. The author had a reason, and the reason is good. The deploy is green because a static read of this file finds zero third-party imports — which is a true statement about the text and a false statement about the program.
The fix is not to move the imports back up. That would undo the reason they are there. The fix is to say out loud what the scanner cannot see, using the explicit requirements block:
# requirements:
# oracledb
# openpyxl
def main(term: str) -> dict:
import oracledb # still lazy, still runs on a laptop
import openpyxl
...
The block is the declaration card. It tells the resolver what to pack regardless of where the imports live in the file, so the lazy pattern keeps every advantage it had and stops costing you a runtime death. See #requirements for how that block is written and why pinning versions matters.
One practical note if a build step generates your deployable file: have that step prepend the block, rather than relying on somebody remembering. A rule that depends on memory is a rule that works until the week somebody is busy.
1. A green deploy is a statement about what was checked, not about what will run. The resolver did its job perfectly. It reported on the imports it could see, and there were none. "It deployed clean" and "it has what it needs" are different claims, and only the first one was ever tested.
2. The correct instinct causes the bug. Import-inside-function is standard good practice in plenty of contexts, and it is the only way to keep one file running in two environments with different packages installed. You do not get here by being careless. You get here by being thoughtful in a way the tooling cannot observe.
3. Static and dynamic disagree, and only one of them gets a turn early. This is the same family as There Is No File Here: something reads your code in a way that is not execution, and your assumptions about execution do not survive the translation. Whenever a tool inspects rather than runs, ask what it can possibly know from text alone.
4. The failure is delayed, conditional, and therefore intermittent. If the lazy import sits inside a branch that only runs sometimes, the job succeeds for days before dying. Now the evidence is even worse: the same code, the same deploy, working and then not. That looks like an environment problem, and people will go looking at the worker.
5. Declaring costs nothing and buys certainty. There is no penalty for listing a package in the requirements block that you also import at the top. When in doubt, declare it. The block is cheap, explicit, and readable by a human who is trying to understand what this script actually needs — which is a second, quieter benefit of writing it down.
Windmill builds the dependency list by reading the imports at the top of your file — an import hidden inside a function is a package nobody packs, so if you import lazily you must declare explicitly.
Your scheduled job failed at six in the morning. Was it the driver, the password, one table out of seven, or the folder? The job knew. It was never built to say.
A job that has run every Monday for a month fails at six in the morning. The run page shows one red line and a stack trace that ends somewhere deep inside a database driver. Somebody has to read it, guess which of a dozen things went wrong, change one, and wait for the next run to find out if the guess was right.
Every one of those dozen things was checkable in advance, by a script that takes four seconds. That script is called a preflight, and the interesting part is not that it checks things. It is how it reports them.
Old cars had one warning light. It said "check engine", and that was the whole message. The engine could be low on oil, overheating, or have a loose cap on the fuel tank. The light was identical for all three, so the only way to learn which was to take the car to someone who would plug in a machine and ask.
Modern dashboards have a row of separate lights: oil, battery, temperature, tyre pressure. When one comes on, you know what it is before you stop the car. Nothing about the car became more reliable. What changed is that each part got its own way to complain.
A stack trace is the "check engine" light. A preflight is the row of lights.
A preflight is a small, separate Windmill script that lives next to the real job and checks every layer the job depends on, in order, printing one line per check:
Two rules shape the whole thing. It runs on the worker, with the job's own resources, never on your laptop with yours (The Code Works, the Network Doesn't is the reason). And it reads from the source but never changes it; its only write is the probe file, which it removes.
The result is a dictionary with ok, a list of what failed, and every check
with its detail. Green means the job's whole world is in place. Red names the
piece that is not.
Think of the checks as two columns, not one list.
The database column is a chain: you cannot read a table without a connection, and you cannot connect without a resource. When a link breaks, stopping there is correct, because every later check would fail for the same reason and add only noise.
The destination column is a second chain that does not depend on the first at all. Whether the folder accepts a write has nothing to do with whether the database password expired last night. So a red line in the first column must not stop the second column from running. If it does, you fix the password, rerun, and only then discover the folder problem that was waiting behind it: two failed mornings for two problems that one run could have named together.
The skeleton is short. A record helper prints one aligned line per check,
every check is wrapped on its own, and each column runs whether or not the
other one passed.
# requirements:
# wmill
# psycopg2-binary
import datetime
import wmill
TABLES = ["enrollments", "people", "emails", "phones"]
def main(db_resource: str = "f/resources/db_main",
share_resource: str = "f/resources/share_reports",
folder: str = "weekly") -> dict:
checks = []
def record(label, ok, detail=""):
checks.append({"check": label, "ok": ok, "detail": detail})
print(f"{label:<42} {'OK' if ok else 'FAIL'} {detail}")
# Column 1: the source. Each table on its own line.
db = wmill.get_resource(db_resource, none_if_undefined=True)
record(f"resource {db_resource}", db is not None, "" if db else "not found")
if db:
try:
import psycopg2
con = psycopg2.connect(host=db["host"], port=db["port"],
dbname=db["dbname"], user=db["user"],
password=db["password"])
with con, con.cursor() as cur:
for table in TABLES:
try:
cur.execute(f"SELECT 1 FROM {table} LIMIT 1")
record(f" read {table}", True)
except Exception as exc:
con.rollback()
record(f" read {table}", False, str(exc).splitlines()[0])
except Exception as exc:
record("connect", False, str(exc).splitlines()[0])
# Column 2: the destination. Runs even if column 1 failed.
cfg = wmill.get_resource(share_resource, none_if_undefined=True)
record(f"resource {share_resource}", cfg is not None, "" if cfg else "not found")
if cfg:
from f.lib import share # your own small helper for the file share
probe = f"_preflight_{datetime.datetime.now():%Y%m%d_%H%M%S}.txt"
steps = [
("list the share root", lambda: share.names(cfg, "/")),
(f"list the folder {folder}", lambda: share.names(cfg, folder)),
(f"write, read back, compare in {folder}",
lambda: share.put(cfg, folder, probe, b"preflight probe\n")),
("delete the probe", lambda: share.remove(cfg, folder, probe)),
]
for label, step in steps:
try:
step()
record(label, True)
except Exception as exc:
record(label, False, str(exc).splitlines()[0])
break # these steps DO depend on each other
failed = [c["check"] for c in checks if not c["ok"]]
return {"ok": not failed, "failed": failed, "checks": checks}
The output, on a bad morning, reads like a dashboard rather than a mystery:
resource f/resources/db_main OK
read enrollments OK
read people OK
read emails FAIL permission denied for table emails
read phones OK
resource f/resources/share_reports OK
list the share root OK
list the folder weekly OK
write, read back, compare in weekly OK
delete the probe OK
One line to fix, and it says what to ask for.
1. "Connected" proves nothing about the seventh table. Permissions are granted per object. A login that connects and reads six tables can still be refused on the seventh, and a single missing grant is one of the most common real failures. Check every table the query names, separately, or the check that matters is the one you skipped.
2. A green check on a neighbour is a false green. Listing the share root proves the root. The job writes into one folder below it, and that folder can be missing, misspelled, or readable but not writable for the job's account. Probe the exact path the job uses, and make the probe a real write, because listing a folder and writing to it are two different permissions.
3. Stopping at the first red hides everything after it. It feels tidy to bail out on the first failure. In one real migration an expired database password stopped the preflight before it ever reached the share, so that morning the share's state was simply unknown: it might have been fine, or it might have been the next failure waiting in line. That is why the two columns were made independent. Independent layers report independently. Only steps that truly depend on each other stop the chain.
4. A preflight that can crash is not a preflight. If the driver import is
at the top of the file and the driver is missing, the script dies before it
prints a single line, which is the stack trace you were trying to avoid. Import
inside the check so a missing library becomes a red line. Then list those
libraries in the # requirements: block, because lazy imports are invisible to
the dependency scan (Green Deploy, Runtime Death).
5. It is not a one-time setup step. Passwords expire, grants get revoked, folders get reorganised, resources get edited. Keep the preflight next to the job, permanently, and run it after every change to anything the job touches. Four seconds against a failed Monday is not a close call.
A preflight runs where the job runs and asks every layer its own question, so a failure arrives as the name of the thing to fix instead of a stack trace to read.
Your dry run passed. It never touched the one part that breaks: the write.
You built a job that produces a file and drops it into a shared folder. Before you trust it with a schedule, you test it the careful way: a dry run that runs the query, builds the file, and writes nothing. It passes. You schedule it.
The first scheduled run fails, on the write. Of course it does. The query was the part you had already run a hundred times while developing; the write was the part that had never run at all. The dry run tested everything except the thing that was new.
A band rehearsing in a garage proves they can play the songs. It proves nothing about the concert hall: whether the power on stage trips under their amps, whether the hall's speakers buzz, whether the monitors reach the drummer.
So before the doors open, the band does a sound check. Real instruments, real cables, the hall's own speakers, at full volume, in the room where the concert will happen. The only thing missing is the audience. When it is over the hall looks exactly as it did before, but every piece of the real path has carried a real signal.
A dry run is the garage. A test run is the sound check.
A job that publishes something has three useful modes, and the difference between them is how far down the real path each one goes:
_TEST_report.xlsx, read it back, compare it
byte for byte with what was sent, delete it, and confirm it is gone.The point is how small the gap between the last two is: one name and one delete. A test that takes a different path from the real run is testing a different path. A test run that differs only in the name has exercised the folder, the account's write permission, the network route from the worker, and the file library, which is everything the dry run skipped.
In Windmill the switch is just a parameter. test_run: bool = False becomes a
toggle on the run form (Shaping the Form), so anyone can run the test
from the form without touching code.
Read the lanes left to right. All three start the same way. The dry lane stops before anything leaves the worker. The test lane goes all the way to the share and back, then cleans up after itself. The real lane is the test lane with the cleanup replaced by "keep".
The check that does the heavy lifting is the read back. After the write, the
job opens the file it just wrote, reads every byte, and compares. Only a match
counts as delivered. The result reports what the share holds, not what the job
hoped to send: the name, the size, a checksum, and for a test run,
deleted_after_test: true.
import hashlib
def main(test_run: bool = False, dry_run: bool = False) -> dict:
blob = build_the_file() # query + workbook, same in every mode
if dry_run:
return {"mode": "dry", "bytes": len(blob)}
name = "weekly_report.xlsx"
if test_run:
name = "_TEST_" + name
got = put_and_verify(cfg, folder, name, blob)
if test_run:
remove(cfg, folder, name)
got["deleted_after_test"] = name not in list_names(cfg, folder)
return {"mode": "test" if test_run else "real", **got}
def put_and_verify(cfg, folder, name, blob) -> dict:
"""Write, read back, compare. Raise if the share holds anything else."""
write(cfg, folder, name, blob)
back = read(cfg, folder, name)
if back != blob:
raise RuntimeError(f"{folder}/{name} on the share differs from what was sent")
return {"name": name, "bytes": len(back),
"sha256": hashlib.sha256(back).hexdigest()}
Note the "mode" key in the result. It looks like decoration. The next section
explains why it is not.
1. "No exception" does not mean "written". It is natural to assume a write call that returns quietly has succeeded. One file-share helper in a real pipeline returned a list of per-file results, each with its own success flag, and raised only on connection errors. A refused write came back as a perfectly ordinary return value. Code that did not inspect that list would have reported success for a file that was never there. Reading the file back closes that gap no matter what the library does.
2. Test in the real folder, not a test folder. The folder is the thing
under test: its existence, its path, and whether this account may write there.
A separate test folder proves a different folder. The _TEST_ prefix is what
keeps the real folder safe, and deleting the file and checking it is gone is
what keeps it clean.
3. The toggle's default is off, so the Run button is the real thing. A
boolean defaulted to False is the safe default for a scheduled job, and it
means that opening the form and pressing Run performs a real run. It is very
easy to believe you ran a test when you ran the real job. It has happened: a
run reported as "the test worked" turned out to be the real run, which luckily
was also fine. Put the mode in the result so nobody has to remember which
switch was where.
4. A test run may read a test database; a real run must not. A test run publishes nothing, so pointing it at a test copy of the data is fine and often the only option in a staging workspace. A real run should refuse unless it is reading production (A Resource Name Proves Nothing). The two modes have different rules about the source, not just the name.
5. Run it after the preflight, not instead of it. The preflight (The Preflight That Names the Layer) says every layer is reachable with a tiny probe. The test run says the real file, at its real size, makes the whole trip. A share that accepts a 30-byte probe can still refuse a large workbook, or time out halfway through one.
A test run takes the real path under a throwaway name (write, read back, compare, delete) because a dry run only proves the half that rarely breaks.
They want one file, same name, replaced every week. The obvious way to do that is also the way to leave them nothing at all on a Monday morning.
The first version of a scheduled export usually writes a new file every run, with the date in the name. Within a month the folder holds a dozen of them and someone asks the reasonable question: can it just be one file, always with the same name, replaced each time? Their spreadsheet links to it. Their bookmark points at it. They do not want to hunt for the newest one.
So you change the job to overwrite report.xlsx. It works for weeks. Then one
Monday the job dies halfway through the write, and the person who relies on
that file opens it to find it damaged. Last week's good copy is gone too,
because overwriting it was the first thing the job did.
Think about a notice board in a hallway. There is a careless way to update a notice: take the old one down, walk back to your desk, write the new one, come back and pin it up. For as long as that takes, the board is empty, and if you get called into a meeting on the way, it stays empty all afternoon.
The careful way is the opposite order. Write the new notice first, at your desk. Read it over. Walk to the board with it in your hand, and only then swap it for the old one. At every moment, anyone walking past finds a notice on the board. And if someone is standing in front of the board reading the old notice, you wait; you do not pull it out of their hands.
Overwriting a file in place is, from the reader's point of view, delete followed by write. Opening the old name for writing empties it first, then the new bytes stream in. Anything that interrupts the stream (a dropped connection, a worker restart, a full disk) leaves a partial file under the good name, and the previous version no longer exists anywhere.
The safe pattern adds one step at the front and one at the back:
_new_report.xlsx, in the same folder. Read it back and compare
(The Test Run That Leaves No Trace). The new version is now safe on the share.If step 2 fails, the new version is not lost, and how the old one fares depends on when it failed. The common case, a file that is open and locked, is refused when the job tries to open it, before a single byte is written, so the old file is intact. The rare case, a failure halfway through the stream, can leave a partial file under the real name, exactly as in the hook. The difference is what sits beside it: a verified copy of the new version, and a run that failed loudly and says where that copy is.
On a local disk the classic move is to write a temporary file and rename it over the old one, because a rename there is atomic. On a network share, whether a rename may replace an existing or open file depends on the server and on the library, so this pattern does the simplest thing that is safe everywhere: it writes the verified bytes twice.
The most common failure is not a crash. It is a person. Someone opened
report.xlsx on Friday afternoon and left Excel running over the weekend. On a
Windows file share, a workbook open in Excel is locked, and the share refuses
to let anyone overwrite it.
Walk through it with the safe pattern. Stage: _new_report.xlsx is written and
verified. Replace: refused, the file is locked. The job fails, loudly, and the
folder now holds the old report, still readable by the person who has it open,
and next to it the new verified one. The next run replaces both. Nobody lost
anything, and the error message told them exactly what to close.
STAGING_PREFIX = "_new_"
def replace(cfg, folder: str, name: str, blob: bytes) -> dict:
"""Replace last run's file without ever leaving only a broken copy."""
staging = STAGING_PREFIX + name
existed = name in list_names(cfg, folder)
put_and_verify(cfg, folder, staging, blob) # 1. the new file is safe
try:
got = put_and_verify(cfg, folder, name, blob) # 2. now touch the real one
except Exception as exc:
raise RuntimeError(
f"could not replace {folder}/{name} ({exc}). It is usually open in "
f"Excel. The new, verified file is in {folder}/{staging}; the next "
f"run replaces both.") from exc
remove(cfg, folder, staging) # 3. tidy up
got["replaced_previous"] = existed # 4. report
return got
The first run creates the file, so it reports replaced_previous: false. Every
run after that should report true.
1. A fixed name turns "write a file" into "replace a file people are using." With dated names, a failed run costs one missing file and the old ones remain. With one fixed name, every run is an operation on the copy someone depends on. The job's failure mode changed when the naming did, even though the code barely moved.
2. Someone always has it open. Not at six in the morning, usually, which is one more reason to schedule the job then. But eventually a laptop stays awake over a holiday with the workbook open. The lock is a feature: it is the share protecting the reader. Your job's part is to fail clearly and leave both copies where people can find them.
3. replaced_previous is a tripwire, not trivia. If a weekly run that has
reported true for months suddenly reports false, the job did not find its
own file. Someone moved, renamed or deleted it. That is worth knowing on the day
it happens, not weeks later.
4. One-off runs need their own name. If the same job can also produce a file for a specific period on request, give that run a different fixed name, so a one-off request never overwrites the file the schedule maintains.
5. Switching to a fixed name leaves the old dated files behind, and they are not yours to delete. The job can prove which file it creates: the one with its own name. It cannot prove who else put files in that folder, or who still uses them. Clean-up of the old dated files is a question for the folder's owner, not a line in the job. The same goes for any diagnostic that tidies up by prefix: let it remove only the prefixes your own scripts write, staging prefix included.
Write the new file under a staging name and verify it before you touch the old one, so whatever goes wrong, the folder always holds a good copy and the error says where.
The resource had the right-sounding name. The job ran green. The file had a quarter fewer rows than it should, and nothing anywhere turned red.
You are moving a job onto a new workspace, and you need a database resource. In the resource list there is one whose name contains the word for production and the name of the service account your team uses for jobs like this. It is exactly what you were looking for. You point the job at it.
In the best case, it fails on the first query: the table does not exist. In the worst case, it runs. Every step is green, the file lands in the folder, and it has about a quarter fewer rows than it should. Nobody notices, because a scheduled job has nobody comparing row counts by hand.
In a shared office kitchen there are two glass jars of white crystals. One label says sugar, the other says salt, both in someone's handwriting. Last week somebody refilled the sugar jar from the wrong bag.
Nothing about the jar tells you. The label is still neat. The crystals look the same. A recipe that uses the wrong jar still follows every step correctly, and produces something that looks right and tastes wrong. The only test that works is to put a pinch on your tongue, which takes two seconds and is the one thing nobody does, because the label already answered the question.
A resource (Resources & Resource Types) is a typed JSON object stored at a path: host, port, database, user, password. The path is a name a person chose. Nothing in Windmill checks that the name and the values agree, and there is no reason it could: the platform cannot know what "production" means to you.
That leaves three ordinary ways for a name to mislead:
So the question "which database is this?" has exactly one reliable answer, and it comes from the database.
Picture two wrong resources and what each one does to your job.
The first points at a different system altogether. The query names tables that system has never heard of, and the job stops with "table or view does not exist". This is the friendly failure. A preflight (The Preflight That Names the Layer) catches it in seconds, with the table name in the red line.
The second points at a test copy of the right system. Same schema, same tables, same columns, slightly older and slightly smaller data. Every query succeeds. The job is green. In one real migration that copy returned about three rows for every four that production held, and nothing failed. This is the dangerous one, because there is no error to notice: only a number nobody is looking at.
Ask the database who it is, print the answer, put it in the output, and refuse to publish when the answer is wrong.
import psycopg2
import wmill
EXPECTED_DB = "warehouse" # what production calls itself
# The same question in other dialects:
# Oracle: SELECT sys_context('USERENV','DB_NAME'),
# sys_context('USERENV','SERVICE_NAME') FROM dual
# SQL Server: SELECT @@SERVERNAME, DB_NAME()
def main(db_resource: str = "f/resources/db_main", test_run: bool = False) -> dict:
db = wmill.get_resource(db_resource)
con = psycopg2.connect(host=db["host"], port=db["port"], dbname=db["dbname"],
user=db["user"], password=db["password"])
with con, con.cursor() as cur:
cur.execute("SELECT current_database(), inet_server_addr()")
database, server = cur.fetchone()
print(f"{db_resource} reads {database} on {server}")
if database != EXPECTED_DB and not test_run:
raise RuntimeError(
f"{db_resource} reads {database}, not {EXPECTED_DB}. "
"Refusing to publish from it.")
rows = run_the_query(con)
# ...and write `database` and `server` into the file itself, on a small
# "About this file" sheet, next to the time it was generated.
return {"database": database, "server": server, "rows": len(rows)}
A test run may read a test copy; that is what it is for (The Test Run That Leaves No Trace). A real run may not.
1. The loud failure is the kind one. It is tempting to think the worst outcome is a job that crashes. The worst outcome is a job that succeeds against the wrong data. Design for the quiet case: the refusal in the code above exists precisely because nothing else would have stopped it.
2. Same path, different workspace, different value. A job promoted from a staging workspace to the production one keeps its paths and silently changes its plugs. That is the feature that makes staging useful. It is also why you should never infer the target from the path, and why the check belongs in the job rather than in your memory of how the resource was set up.
3. Even reading the address can mislead. Resource shapes vary. One Oracle
resource in a real workspace kept the whole connect string in its database
field and had no host at all; code that assembled host:port/service printed
None:1521/ followed by the entire connect string, and looked broken when it
was not. Print the address from the fields that exist, and then ask the
database anyway.
4. Row counts are not a check unless something compares them. "We would notice if the numbers dropped" assumes someone is watching. On a schedule, nobody is. The comparison has to live in the job: a refusal on the wrong database, and a refusal on zero rows where zero is not plausible.
5. Write the answer into the output. A small "About this file" sheet with the database name, the generation time and the row count costs nothing. The person who opens the file can check where it came from without asking you, and the day something looks off, the first question is already answered.
A resource path is a label a person wrote; before a job publishes anything, ask the database who it is and refuse if the answer is not the one you expect.