Are you spending a fortune on tokens? How does $0.05 per million sound? This morning, I tried one of the cheapest LLM’s in existence. Turns out, it’s surprisingly powerful.
Llama 3.1 8B on Groq costs just $0.05 per million input tokens and $0.08 per million output tokens. That’s over 100x cheaper than frontier models.
I was shocked at how capable this little LLM is! From basic coding tasks to parsing documents, Llama crushed one task after another.
Let me show you what this bargain basement model can do…
Round #1: Making a Basic Landing Page

For our first round, I had Llama make me a basic landing page. I told it to give me a simple MS-DOS format.
Llama gave me a nice, simple page in a jiffy. I’m impressed that such a cheap model can handle coding tasks this well!
The page I made with Grok Build was a little prettier. But if appearance is less of a concern than cost, Llama works nicely.
I’m giving this round an A-.
Round #2: Extracting Details from a Document

Lots of founders use AI to extract info from documents. This is a repetitive task where a cheaper model could really help.
Can Llama extract info reliably?
I gave Llama my blog post on Digital Ocean’s seed round. I asked it to tell me the amount of money they raised in their funding round: $3.2 million.
Llama nailed it! I’m giving this round an A.
Round #3: Summarizing a Document

Another common task for AI is summarizing documents. On high-volume tasks like this, token costs can really add up.
Can Llama help?
I fed Llama an article from The Japan Times about surging sales of portable air conditioners. Llama gave me an excellent summary, explaining that low prices and easy installation have made portable AC’s very popular in Japan.
Llama produced this summary incredibly quickly: just 256 milliseconds. Llama gets an A+ on this round!
Wrap-Up
Llama 3.1 8B earns an A overall in my testing.
This little model really impressed me. On task after task, it gave me great outputs with incredible speed. I’ve gotten worse results from more expensive models!
If you’re doing something really complex, like debugging a huge codebase, you’re better off using one of the frontier models. But for many common tasks, Llama is more than sufficient.
If you’re struggling with token costs, try embedding Llama into your application!
More from the blog:
I Tested Mira Murati’s New Open Model. Here’s Where It Wins (and Loses)
How I Built a Slick Landing Page in 28 Seconds with Grok 4.5, Elon’s New Coding Agent
China’s GLM 5.2: The Most Powerful Open-Source Model Yet — But Does It Deliver in Real Life?
Save Money on Stuff I Use:
This platform lets me diversify my real estate investments so I’m not too exposed to any one market. I’ve invested since 2018 with great returns.
More on Fundrise in this post.
If you decide to invest in Fundrise, you can use this link to get $100 in free bonus shares!
I used this app every single day to dictate to my computer, I’m even dictating this text using Wispr Flow! It’s way better than Apple’s native dictation.
My productivity is up about 25% since I started dictating rather than typing. I’m also less tired and stressed.
Get a free month of Wispr Flow Pro here!
Leave a comment