Enhance document processing in scrapers

- Updated pediatric multiple sclerosis article reference ID.
- Improved document_to_markdown.py to handle differential diagnoses, tables, anatomy, and cases more effectively.
- Added logging for cached entries of DDX, Tables, Anatomy, and Cases.
- Adjusted file pattern matching to process all JSON files by default.
- Fixed document ID extraction logic for various data types.
This commit is contained in:
Ross
2025-10-18 15:41:33 +01:00
parent 5983ca3252
commit 66512aa439
71 changed files with 1176 additions and 15 deletions