š¤ Monthly Metal Crawler and local LLMs
How a Swift and MLX crawler became a practical experiment in using different local models for extraction, comparison and recommendation.

For a while, running an LLM locally was technically impressive, but difficult to turn into something I actually wanted to build around.
I tried lightweight models as an experiment a few times. A 4B model could run quickly on a laptop, respond almost immediately, and was surprisingly capable at simple tasks. But once I asked it to work with real-world text, preserve context across several sections, or reliably return structured data, the limitations became obvious.
At the other end were the genuinely powerful models. Those were designed to run on full-scale servers and simply couldn't fit into normal home hardware.
More recently, this started to change. With enough unified memory, especially on Apple Silicon, it became possible to run models far larger than what would fit on a conventional laptop GPU. Apple's MLX framework is explicitly designed around this advantage: CPU and GPU operate on the same unified memory instead of constantly copying tensors between separate memory pools. That architecture makes running unusually large models on high-memory Macs much more practical. (MLX unified memory documentation)
But what about something more powerful that still fits into home hardware?
This was never an empty part of the model catalogue. Google released Gemma 3 in sizes up to 27B and later published INT4 versions that made the 27B model practical on a 24 GB consumer GPU. Mistral had 24B-class models. Qwen already offered 14B, 30B and 32B variants in 2025. (Google Developers Blog)
Still, from the perspective of actually building something on a laptop, there was a very practical gap.
The small models were fast enough to put almost everywhere, but not always reliable enough. The very large models were capable enough, but expensive in memory and, more importantly, slow enough that putting them into every step of a workflow completely destroyed the development loop.
And now, the 20ā40B range
A lot of this progress is currently coming from Chinese model developers.
Alibaba's Qwen family is probably the clearest example. During 2026 it released Qwen3.5 and Qwen3.6 models at 27B, 35B-A3B and larger sizes, followed in August by Qwen3.8-27B. The latest 27B model is dense, supports a 262K context and allows configurable reasoning effort. It is also released under Apache 2.0. (Qwen3.8)
DeepSeek and Z.ai are pushing the same idea much further. DeepSeek V4 Flash has 284B total parameters but 13B active, while GLM-5.3-Flash has 320B total and 18B active. Of course, those models still require a lot of memory because the weights need to live somewhere. But they show where the architecture is going: significantly more capability without proportionally increasing the amount of computation required for every generated token. (DeepSeek V4 preview)
And this is not exclusively a Chinese-model phenomenon. Mistral Small 4, released in March 2026, is another attempt to combine reasoning, multimodality and agentic capabilities in a much more compute-efficient open model. (Mistral Small 4)
Looking at all of this, I thought: this is finally a fantastic use case for my top-spec MacBook Pro.
The only question was which project to try first.
Obviously, in the future, all of them. I have huge plans.
But for now, I decided to tackle the most obvious one for me.
š¤ Monthly Metal Crawler
I listen to and discover a lot of metal albums every month. I also collect CDs and use a high-end digital source to play music files, so CD and hi-res availability matters to me.

There are a few sources I trust, like BangerTV on YouTube and several Instagram profiles. However, I don't have time to check all of them regularly while also filtering out the sub-genres I am not interested in.
Another component is Metal Archives, which is basically a very good metal wiki. Bandcamp is the best source for complete releases where I can actually buy the files or CDs. BangerTV and other editorial sources are useful for discovery, while Instagram can surface smaller releases that might not appear elsewhere.
So, I wanted to automate all of these pieces and get something like this:

That became Monthly Metal Crawler, a macOS/Swift/MLX project that combines a small number of discovery sources, extracts album candidates, enriches them with structured information and finally generates a ranked recommendation report.
The current pipeline looks roughly like this:

An important part of the project was figuring out where not to use an LLM š
Most of the crawler does not need AI
Metal Archives is structured: release date, type, label, genre, lyrical themes, review scores, years active, members, discography information and cover images can all be extracted reasonably reliably from the page structure. So I parse it with code.

Bandcamp is less consistent, but it is still structured enough that normal parsing works better. I ran into a good example while detecting whether an album was available on CD.

My first implementation looked at the whole page for the phrase sold out. That failed.
An album could have sold-out vinyl while the CD was still available. The correct solution was not to ask a language model to understand the page. It was to find Bandcamp's individual purchase blocks and inspect the block corresponding to the CD.
Again, an ordinary HTML parser worked better in this case.
The LLM became useful somewhere else
BangerTV descriptions, review-video descriptions and Instagram captions are human-written text. Their format changes constantly.
One video description might contain five releases. Another might contain one reviewed album followed by several shout-outs. An Instagram caption can mix album names, songs, formats, store information and unrelated commentary.
Writing deterministic parsers for all of those formats would be possible, but it would also mean constantly adapting them to slightly different descriptions, captions and layouts. This is exactly the kind of problem where using a local model started to make sense.
Source extraction became a simple loop: find the album candidates and return them in strict JSON. If the response does not match the schema, the agent sends it back for a second iteration together with the validation errors.
That is also where I deliberately limited the model's responsibility. It does not decide what to crawl, which source to use next, or what to do with the result. All of that is still controlled by the agent. The LLM is just one step in the pipeline, used to convert inconsistent human-written text into structured data.

This ended up being much more reliable than continuously expanding the original prompt and hoping that eventually it would become impossible for the model to misunderstand it.
This is also where experimenting with different model sizes became useful.
I started with a 4B model. For development it was great: it started quickly, responses were almost immediate, and it was good enough to prove that the extraction flow worked.
On cleaner source documents it could also be surprisingly accurate. The problems started with messier text. It could miss a section, lose the relationship between parts of the description, or interpret a casual mention as an actual album recommendation.
JSON itself was not really the problem. The model could follow the schema reasonably well. The difficult part was understanding the source correctly before producing it.
So I moved up and tested a 35B model. The extraction quality improved noticeably. It handled longer descriptions and ambiguous sections better and followed the schema more consistently.
It was absolutely possible to run on my MacBook Pro with 128 GB of memory. It just wasn't something I wanted sitting inside a loop that might execute repeatedly while I was developing and testing the crawler.
Eventually I settled on a much more useful split:
14B ā source extraction
120B ā final recommendationThat decision turned out to be more useful than trying to find one ābestā model for the whole pipeline.
The extraction and recommendation steps might look similar from the outside: both receive some input and return structured JSON. In practice, though, they need very different behaviour from the model.
For extraction, I want the model to be quite boring. It should preserve names, understand the relationships in the source text, and follow the schema consistently. There is very little room for creativity here, and speed matters because this step runs many times during a crawl.
Recommendation is a different problem. By the time the crawler gets there, most of the deterministic work has already been done. The candidates have been collected, deduplicated, validated and enriched with additional metadata. At this point I am no longer asking the model to extract facts from messy text. I am asking it to compare the candidates and make a judgement based on the information already collected.
Each candidate may already contain:
- band and album
- release type, date and label
- source evidence
- genre and lyrical themes
- review count and score
- years active, member count and full-length album count
- Bandcamp URL, digital formats, hi-res and CD availability
- cover image and enrichment errors
At the recommendation stage the problem becomes much more comparative. I want the model to look across all of the collected information and decide which albums are actually interesting, which signals support that recommendation, whether there are any reasons to be cautious, and ultimately whether something looks worth buying rather than just streaming once.
This is where using the 120B model makes much more sense. It only runs after the candidate set has already been collected, reduced, validated and enriched, so I am not paying that cost throughout the whole pipeline.
The whole flow still runs as a single pipeline. One command crawls the sources, extracts candidates, enriches them, runs the final recommendation step and produces the report:
swift run metal-crawler crawl-metal --month 2026-06 --recommendInternally, though, the pipeline uses two different local models for two different jobs. The extraction model runs on one local endpoint, while the larger recommendation model runs on another:
source extraction: http://127.0.0.1:8082/v1
recommendation: http://127.0.0.1:8083/v1For me this was a useful setup because I could tune or replace one model without changing the other. The crawler itself does not care that different models are involved: from the outside it is still one command that produces the complete result.
Generating the final report
One thing I wanted from the beginning was for the final result to be something I could actually use, rather than another chat interface. Glorious HTML! It gives much more freedom to present the results nicely and is simply more interesting to look at than another .md file.
The recommendation model returns structured data with the ranking, reasons for each recommendation, possible concerns, purchase notes, confidence and supporting evidence. The agent then validates that every recommended album actually exists in the candidate set, so the model cannot simply add another album because it happens to know about it.
After that, the agent renders the HTML report. The recommendation model does not generate the HTML itself. Links, cover images, escaping, layout and final artifact generation are all handled by normal code inside the agent.
The crawler currently produces three files:
monthly-metal-recommendation-context.json
monthly-metal-recommendations.json
monthly-metal-recommendations.htmlThe first contains exactly the context sent to the recommendation model. The second contains its structured response. The HTML file is the actual end result I want to open and use.

Keeping these stages separate also made debugging much easier. If I disagree with a recommendation, I can inspect the context the model received. If an album is missing, I can trace whether it disappeared during extraction, deduplication or enrichment. If Bandcamp availability is wrong, I can debug that parser independently without touching the LLM parts at all.
The best part isāthe model is not wrapping the whole application in some opaque decision layer. It is used in a few specific places, while the agent keeps control of the workflow and all inputs, outputs and intermediate results remain visible and testable.
What I learned from it
The project actually started in a very different place.
The early version was a generic local-agent harness with reasoners, tool calls, execution steps, state and storage. Over time, I removed most of that as the actual use case became clearer.
The project went through roughly three stages:
generic local agent harness
ā
monthly metal crawler
ā
deterministic local-LLM research pipelineThat evolution is probably the most useful result of the whole experiment.
When I first started playing with local models, my intention was to give them more responsibility: tools, loops, decisions and orchestration. Once I started building an actual workflow, I ended up moving in almost the opposite direction.
Most of the pipeline is deliberately normal software. HTTP requests, parsing structured websites, deduplication, validation and artifact generation are all deterministic code. There is no real benefit in asking an LLM to do something that can be implemented reliably without one.
The smaller model is useful at the point where the input becomes inconsistent human-written language and I need to turn it into structured data. The larger model comes much later, when all of that data has already been collected and validated and the problem becomes more about comparison and judgement.
For this project, that division worked much better than trying to make the LLM responsible for the whole workflow. The agent controls the pipeline, while the models are used for specific tasks where they actually add something.
Being able to run an LLM locally is not new by itself. We have been able to do that for a whileāwell, two or three years; time flies. What is changing is the quality of the models that fit into this practical middle.
A good 20ā40B model is small enough to become a normal part of the software stack on a capable laptop, while becoming good enough that its role does not have to be limited to simple summarization or experiments.
For me, that opens a much more interesting class of applications. Not simply āChatGPT, but offlineā, but regular software where a few difficult boundariesāunstructured text, classification, comparison, ranking or interpretationācan be handled by a local model without giving up control of the rest of the system.