Configuring a Versatile LLM Harness & Scraping the Web With Scrapy

July 24
58 mins

View Transcript

Episode Description

Which is more important, the model or the “harness” around an LLM? What are ways to assemble an efficient agentic developer workflow? This week on the show, Ayan Pahwa joins us to discuss harnessing, web scraping, and self-hosting Python applications.

Ayan is a developer advocate at Zyte and an experienced project builder. We discuss a recent article he wrote about creating an extension for the web scraping tool Scrapy. He also digs into his self-hosting setup for Python applications and tools.

Our discussion extends to the complexities of developing effective harnesses. Ayan shares his setup and how he navigated shifting from prompt engineering to context and loop engineering.

This episode is sponsored by HydraDB.

Course Spotlight: Introduction to Web Scraping With Python

In this video course, you’ll learn all about web scraping in Python. You’ll see how to parse data from websites and interact with HTML forms using tools such as Beautiful Soup and MechanicalSoup.

Topics:

  • 00:00:00 – Introduction
  • 00:02:01 – Scrapy and building an extension
  • 00:08:46 – Zyte and the web scraping API
  • 00:11:19 – Sponsor: HydraDB
  • 00:12:22 – noalgotube project
  • 00:15:49 – Homelab & self hosting projects
  • 00:22:19 – What goes into a harness?
  • 00:32:33 – Where did you start exploring LLM tools?
  • 00:36:13 – Local models & edge computing
  • 00:39:31 – ExtractPod and discussing Apple’s AI
  • 00:43:35 – Video Course Spotlight
  • 00:44:54 – Managing token use and tools
  • 00:52:37 – What are you excited about in the world of Python?
  • 00:55:10 – What do you want to learn next?
  • 00:56:38 – What is the best way to follow your work online?
  • 00:56:58 – Thanks and goodbye

Show Links:

Level up your Python skills with our expert-led courses:

Support the podcast & join our community of Pythonistas

See all episodes