Back to feed

Running Uncensored Models on a Free 32 GB Kaggle Server

Kaggle's free notebook with two T4s totaling 32 GB, combined with Ollama and a Cloudflare tunnel, becomes the lowest cost way to run uncensored models through Hermes Agent without owning local VRAM.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — EkA4pqXgta0
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The promise is simple and practical: instead of buying a bigger graphics card, you borrow a free Kaggle notebook that combines two T4 accelerators into 32 GB of memory and run any uncensored model online, then point your favorite agent harness at it as if it were a local server.

The video opens with a familiar frustration from two previous episodes about unrestricted models; viewers loved the ideas but most machines ran out of VRAM, the install failed, and the demo never started because consumer laptops rarely exceed 12 GB.

Kaggle provides that missing headroom for free: choosing GPU T4 x2 gives you two T4 cards totaling 32 GB, a weekly quota of 30 hours that resets every week, and a hard requirement to verify your phone number before any accelerator becomes available, otherwise you stay on CPU only.

After verification the flow is straightforward: hit Create to open a new notebook, open Settings and set Accelerator to GPU T4 x2, and you share the same 30 hour quota whether you pick P100 or T4 x2, with T4 x2 offering more memory and better mixed precision throughput for large models, a limit that feels more predictable than the 12 hour session cap on Colab free.

The notebook does not look like a classic server; it is a stack of code cells that only run when you prefix the command with an exclamation mark, you execute each cell individually with the Play button, and you add or remove cells with the plus control, switching between Code and Markdown as needed.

Setup begins with tiny system tools such as zstd, curl and wget installed quietly via apt, then the official Ollama install script fetched with curl and piped to sh, and because Kaggle notebooks lack systemd you must start the daemon in the background with a subprocess call or the service will block.

Once installed you run ollama serve and when the log prints Ollama started the service listens on port 11434 on the local address, which is still invisible to any external harness, so a public tunnel is required before an agent can reach it.

Model discovery happens on Hugging Face: filter for Text Generation and the Ollama app, type uncensored in the search box and the most downloaded open variants appear, letting you compare Q4_K_M and Q6_K quantized builds and see a green check when the chosen file fits the 32 GB budget, with some models offering a 16 GB VRAM optimized build.

To fetch the model you run ollama pull with an exclamation prefix inside the notebook; the progress prints line by line, the system tries to run the model once download finishes, and you can stop that attempt immediately because final inference will happen through the tunnel, not inside the notebook.

To expose the local service you install the Cloudflare tunnel binary, unpack it and launch it pointed at localhost on the Ollama port, which yields a public URL that you can open in a browser to confirm Ollama is running, and if the check fails re-running the same cell usually fixes it.

The last leg is the Hermes Agent connection: pick Custom Endpoint in the model menu and enter its number, paste the public URL with /v1 appended as the Base URL, type ollama as the API key, select the pulled model, and if the first call times out retry once after the service warms up and answers the who are you test noticeably faster than the paid cloud demo, with the full command list shared as a TXT in the creator's Discord.

Visualization: nodesdaily AI

VRAM comparison

  • Kaggle T4 x232 GB
  • Kaggle P10016 GB
  • Colab Free16 GB
T4 x2 total memory lead is decisive for larger models.
PlatformGPUVRAMSession / WeeklyPrice
Kaggle T4 x22x T432 GB9 hr / 30 hrFree
Kaggle P100P10016 GB9 hr / 30 hrFree
Colab FreeT416 GB12 hr / variableFree

AI commentary

"In my view this setup is a field-tested bridge for anyone avoiding expensive cloud rentals; it combines free quota, tunneling and open models in one flow and therefore speeds up experimentation."

AI assessment

Steelmanning the other side, I have to say Kaggle notebooks were built for competitions and light experiments, not as a persistent model server. The acceptable use policy explicitly bans server farming and resource abuse unrelated to machine learning. Treating 30 free hours every week as a guaranteed production host is optimistic at best.

The limits are clear to me. A single T4 delivers about 8.1 TFLOPS in half precision and 320 GB/s of bandwidth while an A100 delivers 312 TFLOPS and 2039 GB/s, roughly a 38x gap. Sessions are capped around 9 hours and the environment is ephemeral, so the pulled model is wiped when the notebook stops and must be re-downloaded, while the 2-core CPU often throttles data loading.

On incentives, the Discord TXT creates a small funnel toward the creator community, which is natural. On security, exposing the Ollama endpoint through a public tunnel without authentication is risky given past CVEs, and anyone with the URL can reach the same port. That is why the 30 hour, 9 hour and 16 GB figures shown should be cross-checked against Kaggle docs and the green compatibility check rather than taken at face value.

My practical take is that this method shines for short experiments, Q4 builds in the 7 to 13 billion range, student projects and quick checks of uncensored behavior. It is not a fit for 70 billion models, very long context or always-on production. For those workloads a billed accelerator with persistent storage on Thunder Compute or Runpod provides more predictable uptime than any Cloudflare or ngrok tunnel.

Sources

9 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

hardware · running · uncensored · models · free · kaggle · nodesdaily

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…