Attention — Jul 21, 2026

What caught my attention on Jul 21 (12 items).

· attention

New UK PM Andy Burnham puts AI Britain into the Cabinet

Pivot to AI (Article:18:16) - Published 2026-07-20 18:00 - Discovered 2026-07-21 18:16

In this episode of Pivot to AI, host David Gerard examines UK Prime Minister Andy Burnham’s recent cabinet reshuffle, which fundamentally restructures the nation’s approach to artificial intelligence governance. The most notable change is the dissolution of the Department for Science, Innovation and Technology (DSIT), established in 2023 as Britain’s dedicated AI department. Its functions are bein…

Local LLM performance tuning

NoPlaceLikeLocalhost (Article:20:02) - Published 2026-07-20 18:00 - Discovered 2026-07-21 20:02

This video explores practical performance tuning techniques for local large language models, using the dense Qwen 3.6 27B architecture as a case study. While dense models generally outperform mixture-of-experts (MoE) models in reasoning and capability, they are inherently slower because every parameter must activate for each inference. The host establishes a baseline by loading a Q4 quantized vers…

LLM Battle 2 - Qwen, Ornith, Qwythos, Gemma 4

NoPlaceLikeLocalhost (Article:22:50) - Published 2026-07-19 18:00 - Discovered 2026-07-21 22:50

In this episode of “LLM Battle,” the host challenges four local large language models—Qwen 3.6 27B, Ornith 35B, Qwythos 9B, and Gemma 4 12B—to a coding competition. The qualifying round tasks each model with building a single-file Python clone of the classic arcade game Asteroids using Pygame. The prompt specifies graphical constraints, physics rules, and a creative bonus feature that must inten…

Local AI voice cloning

NoPlaceLikeLocalhost (Article:22:51) - Published 2026-07-16 18:00 - Discovered 2026-07-21 22:51

In this episode, the creator demonstrates how to equip the local AI coding assistant Open Code with full voice interaction capabilities, combining speech-to-text input with AI-generated voice output. The setup begins with open-code-voice, a companion CLI tool that runs a lightweight Whisper model entirely on the CPU. By launching Open Code on a specific network port and piping audio through the…

Automating image tagging with a local LLM

NoPlaceLikeLocalhost (Article:22:53) - Published 2026-07-15 18:00 - Discovered 2026-07-21 22:53

This video demonstrates how to automate the tagging of personal photo collections using a locally hosted multimodal large language model. The creator begins by explaining that standard text-only LLMs cannot process images; users must specifically select a multimodal model and download a separate mmproj (multimodal projector) file to enable visual understanding. After configuring llama-server w…

Running a local LLM with two GPUs

NoPlaceLikeLocalhost (Article:22:54) - Published 2026-07-14 18:00 - Discovered 2026-07-21 22:54

In this video, the creator upgrades a custom-built AI workstation named “Saturn” by adding a second NVIDIA GPU to enable running larger local large language models. Originally equipped with an RTX 4070 (12GB VRAM), the system now incorporates an older RTX Pro 4000 (24GB VRAM), creating a combined 36GB VRAM pool. However, this upgrade introduces a hardware trade-off: the motherboard’s secondary PCI…

Using a local AI for creative writing?

NoPlaceLikeLocalhost (Article:22:55) - Published 2026-07-13 18:00 - Discovered 2026-07-21 22:55

In this video, the host explores whether artificial intelligence can effectively serve as an assistant for creative writing projects, ultimately offering a qualified “yes” while outlining three distinct use cases. The first, and most recommended, application is mechanical editing. By feeding a chapter into an AI, writers can efficiently catch grammatical errors, punctuation mistakes, and regional…

LLM battle 1

NoPlaceLikeLocalhost (Article:22:56) - Published 2026-07-12 18:00 - Discovered 2026-07-21 22:56

In this video, the host of “No place like localhost” conducts a comparative test of several local large language models to determine which can best generate a new full-screen music visualizer for a personal hobby application. Using a consistent prompt requesting floating, gently moving hexagons with a light, airy aesthetic, gradients, transparency, and customizable settings, the host evaluates the…

Building agents and skills with OpenCode

NoPlaceLikeLocalhost (Article:22:58) - Published 2026-07-09 18:00 - Discovered 2026-07-21 22:58

In this episode, the creator continues their series on running a local LLM by demonstrating how to customize OpenCode through custom agents and skills. To understand OpenCode’s default behavior, they first set up a man-in-the-middle proxy to inspect network traffic. This reveals that the default system prompt is excessively long, consumes valuable context space, and enforces a rigid, overly concis…

Setting up context in a local LLM

NoPlaceLikeLocalhost (Article:22:59) - Published 2026-07-07 18:00 - Discovered 2026-07-21 22:59

This video provides a practical guide to configuring context windows for local large language models, framing the context window as the AI’s short-term memory. The presenter stresses that selecting the right size is a critical balancing act: setting it too large can exceed available VRAM, causing severe performance degradation or complete failure, while setting it too small restricts the model’s a…

Building llama.cpp from source

NoPlaceLikeLocalhost (Article:23:00) - Published 2026-07-06 18:00 - Discovered 2026-07-21 23:00

This video demonstrates how to compile and run llama.cpp from source on a Linux system equipped with an NVIDIA GPU, with the ultimate goal of creating a fully local AI coding assistant. The host begins by installing essential development dependencies, including CMake, Git, OpenBLAS, and the latest NVIDIA CUDA toolkit. After cloning the actively maintained llama.cpp repository, he configures an…

Building a completely local AI code assistant

NoPlaceLikeLocalhost (Article:23:02) - Published 2026-07-05 18:00 - Discovered 2026-07-21 23:02

The video addresses the rising costs and token-based billing models of cloud AI coding assistants like GitHub Copilot and Claude Code, advocating for a fully local alternative that eliminates subscription fees, token caps, and rate limits. The host demonstrates how to repurpose an aging 2020 home lab PC into a functional local LLM host by making a single strategic hardware upgrade: swapping a 2GB…