Update Crud functions for new database structure & utlize 10x performance in the catalog data fetchers with new db structure and mongodb indexes. Added support for multi season pack torrents, Updated metadata handling to use a centralized fetching mechanism and improved filters for matching existing entries. Refactored episode processing logic to account for season and episode numbers more flexibly. Enhanced querying pipelines for better performance in catalog and metadata retrieval.
- Updated streaming providers to allow adding torrents by file for webseed and private torrents.
- Implemented functionality to store the content of torrent files.
Introduced a new Jackett scraper to support torrent indexers and refactored Prowlarr scraper for consistency. Enhanced support for advanced individual indexer searching, searching capability with category validation, and healthy indexer management across both scrapers.
This update allows handling expected 404 errors without raising exceptions. Updated all relevant calls in the `bt4g` scraper to use this flag where applicable, improving error handling for requests likely to return 404 responses.
This commit introduces the BT4G scraper to assist in scraping torrent streams for both movies and series. It includes new configuration options, caching mechanisms, and processing logic tailored for the BT4G platform. Additional settings around scraper limits and filtering have been added to ensure robust and efficient integration.
Moved Redis client setup to a new module `redis_database` and updated all references across the codebase. Also added a new configuration setting for redis_max_connections for better resource management.
* Add support for comprehensive scraper metrics tracking
Renamed scrape_and_parse to _scrape_and_parse, adding comprehensive error handling and logging metrics for various scenarios like timeouts, HTTP errors, and validation issues. Incorporated a metrics class in base_scraper.py to summarize statistics related to scraping performance, which is now used across different scrapers including prowlarr, torrentio, and zilean.
* Enhance Prowlarr individual indexer searching logic with indexer healthcheck management & comprehensive metrics summary
Added detailed health checks for indexers and included their statuses in logging. Enhanced background search to use circuit breakers and handle indexers in chunks, improving reliability and fault tolerance.
* Handle custom id items with imdb id based on title and year matching for moving to imdb title
* handle exception onf fetching prowlarr indexer torrent page fetching
* Refactor Circuit Breaker and add state management methods
Reorganized Circuit Breaker class for better clarity and maintainability. Added methods to manage states (`is_closed`, `reset`, `record_failure`, `record_success`) and refactored `call` method to utilize these state checks. Enhanced logging and status reporting with `get_status` method.
* Simplify series metadata & episode data retrieval logic
Streamlined the `get_series_meta` function by removing complex filtering conditions and reducing the aggregation pipeline.
* Refactor CircuitBreaker for web scraping and enhanced recovery
* Refactor ZileanScraper to use parallel requests for searching and filtering streams with new endpoints
* Fix DLHD scraping & enable DLHD without MediaFlow
* Switch to httpx for async HTTP requests
* Implement caching for PikPak token to reduce login error
* Refactor torrent cleanup logic
* Add inactivity monitor extension to close idle spiders
Introduced `InactivityMonitor` to automatically close spiders that remain inactive for a specified time. This extension checks activity at regular intervals and uses configurable settings for check intervals and inactivity timeouts. If no items are scraped within the timeout period, the spider is closed to free resources.
* Handle TypeError in dynamic sorting of streams
* Refactor torrent info scraper to support pre-processing.
Introduced a pre-processing function mapping to handle specific indexer requirements before parsing the HTML. Added a custom pre-processing function for "TheRARBG" to handle URL adjustments, improving modularity and readability in the `get_torrent_info` function.
* Refine error logging and fix hash key in spider
* Fix Prowlarr not stop on max process limit
* #315: Integrate ScrapeOps logging into all scrapers & spiders
Added ScrapeOps logging to Zilean, Torrentio, Prowlarr, and Prowlarr Feed scrapers to enhance request tracking and error handling. Configured ScrapeOps API key in settings and updated Pipfile/Pipfile.lock with scrapeops-python-requests and scrapeops-scrapy dependencies.
* Refactor scraper cache status handler
* verify torrent before parsing on prowlarr & prioritize magnet on badass_torrents
* Add support for provide locally hosted mediaflow proxy public address and reduce the time leg on private ip address checking
* handle RD exception cases
* do not setup scrapeops when api key is none
* Enhanced dynamic sorting of torrent streams
Revised the dynamic_sort_key function to handle different key types more efficiently with match-case. Simplified error handling and improved logging to capture sorting data in the case of exceptions.
* add missing last update date for metadata
* update domain for nowmesports