The first time you open a GPU rental site, the first thing you see is a price table. The same card is listed by several hosts side by side, and some are clearly cheaper than others. That table is the only number in front of you, so you want to click the cheaper one.
But the bill counts every hour the machine is on. That includes the time spent waiting for downloads, the time spent on an install that fails, and the time before a machine disappears mid-run and leaves nothing behind.
The hourly price never showed what bought nothing. On that same job we paid for 4 runs on RunPod, and 2 got no results back: one because the machine vanished, one because of our own mistake, an old script we stopped after 6 minutes. Those 2 cost at least $0.57 of the $4.65, about 12%. The 6-minute run's cost wasn't recorded separately. To see this, add up every run of the same job, idle minutes included, and divide by the runs that got results back.
Our first job on RunPod and our first day on Vast are not in these numbers. We describe them below, in the section on how we worked out the cost.
We used both hosts between 4 and 7 October 2026. All the money we paid adds up to about $10.74, and the table below breaks it down as far as we recorded it.
Part 1What is a pod, and how is it different from serverless?
A pod is a GPU machine you rent for yourself for a while. You start it and work on it like any other computer. RunPod and Vast both bill rental time by the second (ref 1, ref 6). We get into the machine over ssh, a way to type commands on a remote machine from your own computer.
What comes on the pod when it starts is called an image: a prepackaged bundle with the operating system, drivers and the tools your job needs. You pick it on the host's rent page when you start the machine. Once the machine is up it downloads the image first, which is called the image pull, and the machine is billed the whole time it pulls. Pick the right image and you can start working as soon as the pull ends. Pick the wrong one and you install the rest yourself, and every minute of installing is billed.
There are 2 ways to be done with a pod, and they are very different.
- Stop: the machine is switched off and the GPU charge stops, but a kept disk is still charged for storage while the machine is stopped, on RunPod (ref 5) and on Vast (ref 6).
- Delete: the machine goes back to the host. Rental charges end completely, and anything you haven't copied back is gone.
Whether a file survives depends on which disk it sits on. RunPod's docs name 3 kinds (ref 11). The container disk is temporary storage that comes with the machine and is cleared the moment the pod stops. The volume disk belongs to that pod for as long as you rent it: files survive a stop but are deleted when the pod is deleted. A network volume is storage rented separately from any pod, and its files survive both a stop and a delete.
If you rent on RunPod and want to keep results on the machine, keep them on a network volume. The safer path is to copy them off the machine as you go with scp or rsync, two standard command line tools that copy files between machines over the same ssh connection you use to get in. One of our rented machines was deleted mid-job by our own program.
We copy our results to our own server, which stays on and keeps pulling training results back from the rented machine. Wherever this post says "our server", it means that machine. If you have no server, your own computer works too.
Next to the 2 options people usually think of: a pod sits between paying per job and paying once up front.
- Serverless: you don't rent a machine. You send work in pieces to a provider that runs it, and you pay only while work is running. With no work, it scales down to zero on its own. It suits work that arrives as short requests, such as a model answering questions. RunPod and Vast both offer it (ref 2, ref 14).
- Buying a card: one large payment, then unlimited hours.
- Pod: suits long jobs that run for hours and then end, such as training a model. You control the whole machine without buying a card.
A rule of thumb from the two docs: if your work keeps a machine busy for hours at a stretch, rent a pod and stop it when the job ends, because a pod bills for every second it is on (ref 1). But if work arrives in short bursts with idle time in between, serverless fits better (ref 2).
Our job was training a policy, the "brain" that tells a robot how to move each joint. The tool was IsaacLab, NVIDIA's robot simulator, where a robot practices walking in a virtual world thousands of times. To simulate that world IsaacLab needs Vulkan, a 3D graphics interface, and if the machine's graphics driver doesn't support it IsaacLab won't run. One training step for the policy is called an iteration. On RunPod ours took about 3.5 seconds per iteration, so 1,000 iterations took close to an hour. Along the way, the job saves its training progress to a file every so often, called a checkpoint, which you can use or resume training from. At first we saved every 1,000 iterations. After a run where the machine vanished before saving a single file, we switched to every 250 iterations and copied the latest file back every 10 minutes.
We trained 2 policies. The first, Flat, learned to walk on flat ground. The second, Rough, learned to walk on rough ground and climb stairs. Jobs like these run for hours, so a pod fits them best.
Throughout, we tracked money by run. A run is one rental, one paid machine start, from the moment it starts until the machine stops or vanishes. If a run got a checkpoint back to our server, we call it a finished run. It counts as finished even if the job stopped before reaching its plan, as long as a usable checkpoint came back. A finished run means you got usable results back for your money, and whether those results are good is a separate test: our Vast checkpoint was usable but climbed stairs worse than the checkpoint before it.
If your goal is running a language model to answer questions rather than training, we wrote about when renting a GPU to run your own model pays off.
Part 2How are RunPod and Vast different in practice?
RunPod and Vast are both sites for renting GPU machines. Sign up, add credit, choose a card, start a machine. The first difference is who owns the machine.
RunPod has 2 kinds. Secure Cloud runs in data centers with backup power that can be repaired without switching your machine off. Community Cloud is machines from individual card owners that RunPod has vetted (ref 1). We used only Secure.
Vast is a marketplace. Many card owners, from hobbyists to large data centers, list machines at prices they set themselves (ref 3). Renters pick machines from that list one at a time, so what you get depends on the owner.
On both hosts we rented an RTX 4090, a high-end NVIDIA gaming card. On Vast we ran a Vulkan check before training, a quick test that the machine's graphics driver works.
| RunPod (Secure) | Vast | |
|---|---|---|
| Who owns the machine | data centers with backup power (ref 1) | Many owners who list and price their own (ref 3) |
| RTX 4090 per hour | about $0.75, worked back from our bill, startup included | $0.36 to $0.39 listed on the day we rented; we paid $0.39 |
| Could we ssh in | yes, in every Rough run where the machine stayed up to the end; in the one where the machine vanished, ssh stopped answering along with it | on our first day (4 October) we tried 8 machines and got into 1; on the 7 October training run, yes |
| Image our job could use | the only one we found that works is NVIDIA's isaac-lab image; the general image we tried had no Vulkan | the same image, passed the Vulkan check |
| Image pull time | not timed | 16 minutes, billed throughout |
| Why money was wasted | one pod vanished mid-run, cause unknown | first day: 8 machines tried, nothing trained; in the training run, the 16-minute image pull |
What we did not measure in this job. There are 5 things, and for each one there is a single piece of advice we can still give.
- Renting against buying a card: we didn't compare them, so we have no number. If you weigh it yourself, start from cost per finished run on your own job, not the hourly price.
- Image pull time on RunPod: we didn't time it. The only pull we recorded was on Vast: 16 minutes, about $0.10 at $0.39 per hour. Time your first image pull yourself.
- Stop and delete on Vast: we didn't check which kinds of disk keep their files when a Vast machine is stopped or deleted. What this page says about disk charges on a stopped Vast machine comes from Vast's docs (ref 6), not from what we saw. Copy results off the rented machine as you go, so you don't depend on it.
- Images for other tools: we only tried IsaacLab, so we can't say which image another job needs. Read your tool's docs for what it needs from the machine, then do a short run before a long one.
- Picking a Vast machine you can reach: Vast's offer list shows a reliability score, verified status and a datacenter label you can filter on (ref 7, ref 8), but we didn't record these for our 8 machines on day one, didn't test whether filtering helps, and didn't write down how the other 7 failed. On a marketplace, budget for trying more than one machine before one works.
The 2 hourly figures in the table come from different places. Vast's is the listed price. For RunPod we didn't note the price from the site, so we worked it back from the bill. Our first Rough training run that returned anything, run 1 in the bill table below, trained 3,448 iterations at 3.48 seconds each, about 3.3 hours in total, and cost $2.51. That works out to about $0.75 per hour. The bill also includes startup and a short 6-minute run before it, so RunPod's listed price is probably lower. By how much, we don't know.
On RunPod we paid real money before learning that the general image we used, with only model-training tools, had no Vulkan. The machine started, but the job wouldn't run. In the end we had to use NVIDIA's own image.
Vast's 1 of 8 comes from a single day, 4 October 2026. We tried 8 machines that day and could ssh into one. Installing the training tools on that machine failed. The log from that step was lost, so we don't know why, and we stopped using Vast for a while. We weren't fluent at setting up machines yet, so some of the machines we couldn't reach may have failed because of our own settings. The second time, on 7 October, the machine we got worked normally.
On RunPod, the machine stayed up to the end in 3 of the 4 Rough runs, and ssh worked in all 3. In one of them, run 2 in the bill table, the machine vanished after about 42 minutes and stopped answering ssh along with it. Compare that with Vast on day one, where 1 of 8 machines answered. Our first job, training Flat, which included learning to set the machine up, wasn't recorded run by run, so it isn't counted.
Part 3What does the hourly price hide?
The hourly price, as far as we have it, does tell you the direction. In our job it pointed the same way as cost per finished run: Vast was cheaper. How big the gap really is, we can't tell, because RunPod's $0.75 is worked back from our bill and its listed price is probably lower.
What the hourly price didn't show is how much of the bill bought nothing. The Rough job on RunPod cost $4.65 over 4 runs. The pod that vanished took $0.57 and returned nothing. The 6-minute run also returned nothing, but its cost sits inside the $2.51 bill and wasn't recorded separately. So at least $0.57, at least 12% of the $4.65, bought nothing. On Vast the first day cost about $0.84 and trained nothing. Vast's one training run did return a checkpoint, but its first 16 minutes went to pulling the image, about $0.10 with no training. These amounts only show up when you count every minute the machine is on, whether that minute did anything or not, which is what cost per finished run in the Rough job does.
Our wasted minutes, and one near miss, came from 5 places.
- Machines we couldn't reach or couldn't set up. Our first day on Vast cost about $0.84 with no training at all. That figure is mixed with other work on the same account at the time, so it is approximate.
- Image pull time. On Vast it took 16 minutes before the machine was ready to train. At $0.39 per hour that is about $0.10.
- A machine vanishing mid-run. On RunPod a pod disappeared on its own after about 42 minutes. It had trained only 720-odd iterations. At that time we still saved a checkpoint every 1,000 iterations, so nothing was left; we switched to every 250 after this run. That cost $0.57.
- The cost guard misreading a signal. Our cost guard is a program of ours that deletes idle machines. On Vast we had planned to train up to iteration 8,716, and it deleted the machine at 8,000 while the job was still running. That time it cost no money already paid: checkpoint 8,000 was already back, and the 716 planned iterations it cut never ran, so we never paid for them. What we lost was planned work. Had the checkpoint not been copied back, we would have lost the whole run. The section "What to watch for" below has the details.
- Our own scripts breaking. A script is a short program we write to run tasks instead of typing commands one by one. Across this job we started machines about 15 times on the two hosts, and at least 5 of those failed because of our own scripts. For example, one Flat training run finished, but the script couldn't find the output file, so we had to train again. And the very first Rough run we stopped ourselves after 6 minutes because we had run an old version of the script.
If you're renting for the first time, number 5 is the one to watch most, because it follows you to whichever host you move to.
Part 4Every run we paid for
| Host | Job | Runs | What came back | Cost |
|---|---|---|---|---|
| RunPod | First job: Flat training, including learning to set the machine up | several runs, not recorded separately | a working Flat policy, 9,000 iterations | about $4.32 |
| Vast | First day, 4 October | tried 8 machines | nothing trained | about $0.84, mixed with other work |
| RunPod | Rough | 2 runs on one bill: a short run we stopped after 6 minutes, and run 1 | short run: nothing; run 1 finished, 3,448 iterations | $2.51 |
| RunPod | Rough | run 2 | pod vanished after 42 minutes, nothing | $0.57 |
| RunPod | Rough | run 3, resumed from checkpoint 3,447 | finished, iterations 3,447 to 5,716 | $1.57 |
| Vast | Rough | 1 run, resumed from 5,716 | finished, iterations 5,716 to 8,000 | $0.93 |
| Total | about $10.74 |
The last checkpoint of run 1 is numbered 3,447 because iterations count from 0, so run 3 starts at that number.
By host, RunPod comes to $8.97 and Vast to about $1.77. The roughly 15 machine starts mentioned above are all spread across this table, mostly in the first 2 rows, which were not recorded run by run. The Flat run we had to retrain and the first image without Vulkan are also inside the $4.32 row.
Part 5How to work out cost per finished run
Add up every run of the same job on the same host, including the runs that returned nothing, then divide by the number of finished runs. We used the Rough job because it is the only job we trained on both hosts.
| RunPod | Vast | |
|---|---|---|
| Paid runs in the Rough job | 4 | 1 |
| Finished runs | 2 | 1 |
| All runs added up | $2.51 + $0.57 + $1.57 = $4.65 | $0.93 |
| Cost per finished run | $4.65 ÷ 2 ≈ $2.33 | $0.93 ÷ 1 = $0.93 |
RunPod's 4 runs are the 6-minute short run, run 1, run 2 and run 3 from the bill table. The short run and run 1 share the same $2.51 bill, so the short run's cost can't be separated out.
Vast has only one finished run to count, and it is the run our own cost guard cut short, so the $0.93 is the weakest evidence on this page. It still counts as finished because checkpoint 8,000 reached our server before the machine was deleted. We had planned 3,000 more iterations, up to 8,716, so it fell 716 short.
This number answers the money question only. Money is one factor when choosing a host, and a machine you can reliably reach is another; both are in "What to weigh before choosing" below.
This table leaves out our first job on RunPod and our first day on Vast. The first job on RunPod was training Flat, about $4.32, which included learning to set the machine up and got us a working policy. The first day on Vast cost about $0.84: we tried 8 machines, trained nothing, and the cost is mixed with other work on the account. Neither is the Rough job, and they bought very different things, so we don't fold them in.
The runs on the two hosts also differ in length. RunPod's 2 finished runs produced 5,717 iterations in total, and Vast's one run produced 2,284. So we also look at a second measure, cost per 1,000 iterations, to compare work returned against money paid.
| RunPod | Vast | |
|---|---|---|
| Iterations returned | 3,448 + 2,269 = 5,717 | 2,284 |
| Cost per 1,000 iterations | $4.65 ÷ 5.717 ≈ $0.81 | $0.93 ÷ 2.284 ≈ $0.41 |
In this job Vast was cheaper on both measures. Per 1,000 iterations it cost about half of what RunPod did.
But two finished runs against one is too few to know the exact gap, only which way it leans.
Vast's one run had paid time that trained nothing too: the 16-minute image pull. The 716 iterations the guard cut were never run, so never paid for.
Most of what bought nothing on Vast was the first day, which isn't in these numbers. Another day like that on the same job would move Vast's numbers right away.
Part 6Pros and cons of each host
Part 7What to weigh before choosing
Each factor says which way our job pointed.
- Budget: leans Vast. Vast was cheaper both per finished run and per 1,000 iterations. Whichever host you pick, set some money aside for the start, while you're still learning. Our first job on RunPod, training Flat, cost about $4.32, including learning to set the machine up, and got us a working policy. Our first day on Vast cost about $0.84: 8 machines tried, nothing trained.
- Reaching the machine: leans RunPod Secure. In the Rough job, ssh worked in every RunPod run where the machine stayed up to the end, and stopped answering in the one where the machine vanished. On Vast you pick machines yourself; on day one, while we were still learning to set machines up, 1 of 8 answered.
- Image size: does not depend on the host; pull time we couldn't compare. Size belongs to the image, and we used the same image on both. Pull time can also depend on the host's network. We timed it only on Vast, where it took 16 minutes, so time it yourself on your first rental. If your work starts and stops machines often, you may pay for that pull every time.
- Job length: does not depend on the host. On RunPod a pod vanished on its own; on Vast our own cost guard deleted the machine. Either way, the longer the job, the more you lose if the machine goes away near the end. Save checkpoints often and copy them off the rented machine as you go.
- Special requirements of your job: does not depend on the host. Our job needs Vulkan, and the only image we found that works is NVIDIA's isaac-lab image, which works on both hosts.
- Money leaking when you forget to stop a machine: does not depend on the host. A pod bills as long as it's on; forget it for one night and you pay for the whole night. Something should stop machines for you, and it has to read the right signal. Both hosts stop machines on their own when your prepaid credit runs out (ref 4, ref 6). If you have no server of your own, step 5 in "How to rent for the first time" below says what else to use.
Part 8What to watch for: a cost guard that deletes at the wrong moment
Training ran on the rented machine. Our program for stopping money leaks ran on our own server, watching for pods left idle and deleting them. It decided based on the GPU use figure the rental site reports through its API, a channel that lets a program ask the rental site for data directly. If that figure read 0 twice in a row, the program treated the machine as idle and deleted it outright, not just stopped it.
On 7 October 2026, Vast reported that figure as 0.0 twice in a row, half an hour apart, while IsaacLab was training and checkpoints kept arriving on our server about every 11 minutes: at 7,500, at 7,750, then at 8,000. That is 250 iterations every 11 minutes, about 2.6 seconds per iteration on that Vast machine, worked out from the timing; the 3.5 seconds earlier was RunPod's. The program did exactly what it was written to do and deleted the machine at iteration 8,000.
We didn't lose more because we had already copied checkpoint 8,000 back. If we had set it to copy results only at the end of the job, that deletion would have lost every result not yet copied back, which is everything the $0.93 bought. The money itself was spent either way.
The same program had already watched jobs on RunPod without trouble, so we trusted it. Why Vast reported 0.0, we still don't know. The machine was deleted, so there was nothing left to inspect. What we know for sure is that evidence of real work reached our server every 11 minutes, and the program never looked there.
The next part is for readers who have no server of their own: what we took from the two hosts' docs, and a plain fallback. Our own setup, a server that watches the rented machines, comes last. We describe it so the risk is clear, not because a first-time renter needs one.
- As far as the pages we read say, nothing stops an idle machine, but running out of credit does. Neither host describes stopping a machine because it sits idle. Both are prepaid: you add credit first, and when the balance reaches $0 the machines stop on their own (ref 4, ref 6).
- Loading only what you are willing to lose works as a rough cap. Leave automatic top-up off: RunPod's auto-pay and Vast's autobilling both charge a saved card. It is rough because Vast lets the balance go a little below zero. Treat it as a last resort, not the plan, because at $0 a RunPod pod without a network volume is terminated and its data can't be recovered (ref 4, ref 6).
- RunPod's $80 limit won't catch this. By default the whole account can spend at most $80 per hour across all machines, far above one RTX 4090, so a single forgotten machine never reaches it (ref 4).
- A timed stop on RunPod, as an extra only. You type this on the pod, not on your computer: open the web terminal from the pod's Connect button on the RunPod site, or an ssh session into the pod (ref 13). There, type
(sleep 2h; runpodctl pod stop $RUNPOD_POD_ID) &. RunPod sets$RUNPOD_POD_IDinside the pod for you (ref 12). The&at the end runs the whole bracketed command in the background, so you can keep using that terminal while it waits, and after 2 hours the pod should stop itself (ref 5). We haven't run this command ourselves, so we don't know whether it keeps running after you close the terminal, and RunPod's docs don't say. Don't make it your main protection. The stop clears the container disk, so copy your results out before the time is up. The Vast pages we read don't describe anything like it (ref 6, ref 7).
Whichever host you rent from, the protection that doesn't depend on that command is plain: set a phone alarm for the time you plan to stop the machine. When it rings, stop the machine yourself in the host's web console (on RunPod, the console page) and check that it shows as stopped.
Our fix is a lease: a note the job controller (the program on our server that starts and runs training) leaves on our server saying "still working, do not delete", with an expiry time. Before deleting, the guard checks that the note has not expired and that the job controller is still running, both on our own server, and if both hold it only sends us an alert. The note can expire so that a forgotten note cannot keep a machine alive forever.
This fix has passed tests against stand-ins we wrote to replace the real thing, but it has never run against a real, paid pod. If you borrow it, treat it as not yet proven in production. As for whatever auto-shutdown you use now, whether you wrote it yourself or got it elsewhere, check whether it decides on GPU use alone.
Part 9How to rent for the first time
- Run the free checks before starting a paid machine. Two of them use the tools the rental site provides. RunPod has a REST API and a command line program called
runpodctl. Vast has a command line program calledvastai(command line means typing commands into a terminal window). We ran all 4 before renting on RunPod:- Can the image be pulled? Run
docker manifest inspect <image name>, using the name you will type into the host's image field, for usnvcr.io/nvidia/isaac-lab:2.3.2. This needs Docker, the program that pulls and runs images the same way the hosts do, installed on your own computer first. For a public image you don't need to log in. A good answer returns the image's details. If you get "unauthorized" or "not found", fix that before renting. - Is the rental request well formed? Send a rental request with a card name that doesn't exist. With
runpodctlit looks likerunpodctl pod create --image nvcr.io/nvidia/isaac-lab:2.3.2 --gpu-id "NVIDIA GeForce RTX 9999"(ref 9). A good answer rejects the request for the card name only, and no machine starts. If it also trips on another field, your real request would fail there. - What does the system say when a machine is gone? Ask for the status of a machine using an ID that doesn't exist, for example
runpodctl pod get abc123notreal(ref 9). On Vast the matching command isvastai show instance <ID>(ref 10). A good answer is a clear not found. Remember what this answer looks like, so on the day a machine really vanishes, your job knows to stop. - Does your code work with the tools in the image? Open the image's page on the registry it comes from, read which version of the tool it ships, and compare that with the version your training code was written for (for us, IsaacLab in
nvcr.io/nvidia/isaac-lab:2.3.2).
All 4 answered as expected and cost nothing. On Vast,
vastailets you browse listed machines and prices before renting. The first and last checks don't depend on the host. The middle two use each host's own tool, and we haven't run them on Vast. - Can the image be pulled? Run
- Choose an image that fits the job. Each host's rent page has a template or image field where you pick a ready-made image or type an image name. We use NVIDIA's own isaac-lab image on NVIDIA's image registry (nvcr.io). It passed the Vulkan check, so we don't pay for extra installs every time.
- Do a short run first. Start the machine and run something short to confirm it really works before launching a job that runs for hours.
- Save and copy results along the way. Have the job save checkpoints often and copy them back regularly. If the machine vanishes, the job should notice and stop on its own. Ours does this: if ssh fails 5 times in a row, the job asks for the machine's status through the API, and if the answer is not found, it stops immediately. We tested the not-found answer with a machine ID that doesn't exist. It has worked for real once, when our own cost guard deleted the Vast machine: the job noticed the machine was gone and stopped on its own within minutes. We added this stop after run 2, when the RunPod machine vanished, so it wasn't there for that vanish. It has not met a machine that vanished on the host's side since.
- Have something stop machines, but don't let it decide on one number. Something has to stop machines when you forget, and you need to know what it is looking at. If you have no server of your own, use these instead:
- Load only what you are willing to spend, and leave automatic top-up off. Both RunPod and Vast stop machines when your prepaid credit reaches $0.
- Set a phone alarm for the time you plan to stop the machine. When it rings, stop it yourself in the host's web console.
- Schedule a stop from the pod itself, on RunPod, as an extra. We haven't run it ourselves.
The section "What to watch for" above has the details and sources. The thing that stops machines for us is a program we wrote and run on our own server.
- Write down the cost of every run, including runs that returned nothing, then divide by the number of finished runs. That gives you your own cost per finished run. If runs differ in length, also divide by the work returned as a second measure, and keep both for comparing hosts next time.
Part 10A finished run and a good checkpoint are separate questions
Cost per finished run answers only the money question: how much you paid to get something usable back. Which checkpoint to keep needs a separate test, and the host has no say in it.
The goal of Rough was walking on uneven ground, so we tested stair climbing. We simulated it on our own Mac with MuJoCo, a physics simulator, and had each policy walk the robot toward a step 20 times: 5 starting positions, 2 offsets, 2 speeds. A try counts as a pass if the robot climbs the step without falling.
The 3 Rough rows are the same code at different stages of one training job. Checkpoints 3,447 and 5,716 were trained on RunPod; 8,000 continued from 5,716 on Vast. So the table compares training stages, not hosts.
| Policy | Climbed the step |
|---|---|
| Flat 9,000 | 0 of 20 |
| Rough checkpoint 3,447 | 10 of 20 |
| Rough checkpoint 5,716 | 12 of 20 |
| Rough checkpoint 8,000 | 6 of 20 |
Checkpoint 8,000 from Vast climbed the step only 6 times in 20, while 5,716 managed 12 in 20. In this test, 8,000 climbed less often than 5,716 despite the longer training. The Vast run still counts as a finished run for the money, because a usable checkpoint came back; the checkpoint we kept for stairs is 5,716.
Part 11Limits of this comparison
All of this comes from one job: training a robot policy with IsaacLab on an RTX 4090 between 4 and 7 October 2026. In the Rough job, RunPod had 2 finished runs and Vast had one, and the guard deleting a machine by mistake happened once.
There is a lot we didn't record, such as which region the machines were in and RunPod's listed hourly price. Vast's first-day cost is also mixed with other work on the account. Vast prices are set by owners, so they keep moving. If you train a different job or use a different card, your numbers will likely come out differently.
Next time we open a GPU price table, we'll still read the hourly price, because this time it pointed the same way as cost per finished run, even with RunPod's rate worked back from the bill. But we'll ask one more question: how many paid machine starts will it take to get the first finished run? This time, on Vast, we tried 8 machines on day one and got our first finished run 3 days later.
FAQ
Q: Renting a GPU for the first time, should I pick RunPod or Vast?
A: On the one training job we ran on both hosts, Vast was cheaper: $0.93 per finished run against about $2.33 on RunPod, and about $0.41 per 1,000 iterations against $0.81. But Vast had only one finished run to count, and on day one we reached 1 of 8 machines (some of those failures may have been our own setup). If reaching the machine matters most to you, RunPod's Secure Cloud is the tier to look at. In the Rough job the machine stayed up to the end in 3 of 4 runs and answered ssh in all 3, but in our runs one RunPod machine still vanished. Both points come from a single job of ours, so record the costs of your own job and compare again.
Q: What's the difference between RunPod Secure and Community?
A: According to RunPod's docs, Secure Cloud runs in data centers with backup power that can be repaired without switching your machine off, while Community Cloud is machines from individual card owners that RunPod has vetted. We used only Secure, so we can't compare them from experience.
Q: Why not use serverless instead of a pod?
A: Serverless suits work that arrives as short requests and then ends. Training runs for hours on end and needs control of the whole machine, so a pod fits it better.
Q: How do I avoid forgetting to stop a machine?
A: Without a server of your own, load only the credit you are willing to spend and leave automatic top-up off, because both hosts stop machines when the balance reaches $0 (ref 4, ref 6). On RunPod, a pod without a network volume loses its files at that point. On RunPod you can also schedule a stop from the pod itself (ref 5), though we haven't run that command ourselves. In the pages we read, neither host's docs describe stopping a machine for being idle, so also set a timer on your phone for the time you plan to stop the machine. If you run a program that stops idle machines, have it look at evidence that the job is still running too, such as result files still arriving, or a lease with an expiry written by the job. Don't let it decide on GPU use alone. We had a machine deleted mid-training because that figure read 0.
- RunPod, "Pods overview" (Secure Cloud, Community Cloud, per-second billing, the 3 kinds of disk): https://docs.runpod.io/pods/overview
- RunPod, "Serverless overview" (pay only while work runs, scale to zero when idle): https://docs.runpod.io/serverless/overview
- Vast.ai, "Documentation" (GPU marketplace, owners set their own prices): https://docs.vast.ai/
- RunPod, "Billing" (prepaid credits, pods stopped at $0 balance, and terminated without a network volume with their data unrecoverable, auto-pay, low balance alerts, $80 per hour default spend limit): https://docs.runpod.io/accounts-billing/billing
- RunPod, "Manage Pods" (stop vs terminate, volume disk billed while stopped, scheduling a stop with
runpodctl): https://docs.runpod.io/pods/manage-pods - Vast.ai, "Billing" (per-second billing, prepaid credits, instances stopped at zero balance, disk still billed, grace period, autobilling): https://docs.vast.ai/guides/reference/billing
- Vast.ai, "Finding & Renting Instances" (reliability score, verified and unverified machines, Secure Cloud datacenter label, search filters): https://docs.vast.ai/guides/instances/choosing/find-and-rent
- Vast.ai, "search offers" CLI reference (default filter verified=true, reliability and datacenter fields): https://docs.vast.ai/cli/reference/search-offers
- RunPod, "runpodctl pod" CLI reference (
runpodctl pod create --gpu-id,runpodctl pod get <pod-id>): https://docs.runpod.io/runpodctl/reference/runpodctl-pod - Vast.ai, "show instance" CLI reference (
vastai show instance <ID>): https://docs.vast.ai/cli/reference/show-instance - RunPod, "Storage options" (container disk cleared on stop, volume disk survives stop but is deleted with the pod, network volume exists independently of any pod): https://docs.runpod.io/pods/storage/types
- RunPod, "Environment variables" (
RUNPOD_POD_IDset automatically by RunPod): https://docs.runpod.io/pods/references/environment-variables - RunPod, "Connection options" (Connect, Open Web Terminal, SSH): https://docs.runpod.io/pods/connect-to-a-pod
- Vast.ai, "Serverless" (Vast offers serverless too): https://docs.vast.ai/documentation/serverless