Compare commits

..
24 Commits
Author SHA1 Message Date
Jared DillardandGitHub 368c349d58 Fix image paths to deployed images (#30) 2025-08-03 20:47:38 -07:00
Jared DillardandGitHub c816fb1ea8 Fix issue when source_suffix equals source_link_suffix (#29) 2025-07-31 16:17:02 -07:00
Jared DillardandGitHub 92f810592e Add conda-forge version to index.rst 2025-07-21 22:29:24 -07:00
Jared DillardandGitHub ab0eb1dd29 Add conda-forge version to README.md 2025-07-21 14:57:47 -07:00
Jared DillardandGitHub 57f716b2f9 fix docs path 2025-07-12 21:32:01 -07:00
Jared DillardandGitHub 236822885e configure docs theme 2025-07-12 21:19:19 -07:00
Jared DillardandGitHub 46c2dec254 Use first paragraph as summary by default (#22) 2025-06-22 21:00:40 -07:00
Jared DillardandGitHub b56d93d265 Support source file suffix detection (#21) 2025-06-22 19:28:54 -07:00
Jared Dillard 5db5395889 add downloads badge to docs 2025-05-20 01:38:14 -07:00
Jared Dillard 04b0657dc5 add more badges 2025-05-20 01:35:40 -07:00
Jared Dillard 935a964c7e bump version to 0.2.3 2025-05-20 01:20:36 -07:00
Jared Dillard 51f6c71de3 update changelog 2025-05-20 01:18:44 -07:00
Jared DillardandGitHub 70defd3996 Remove get_and_resolve_toctree method (#19) 2025-05-20 01:08:40 -07:00
Jared DillardandGitHub 9ae05c6c13 Simplify _sources lookup (#18) 2025-05-20 00:50:06 -07:00
Jared Dillard 5581979cac make a href 2025-05-18 21:24:14 -07:00
Jared Dillard f70f1a26ec add soft transfer 2025-05-18 21:21:34 -07:00
Jared DillardandGitHub ed50138ae4 Update README.md 2025-05-18 21:12:16 -07:00
Jared DillardandGitHub 8f4d2c07c6 Update docs and README (#17)
* Move readme content to index.rst

* Clean up project name

* Add advanced configuration
2025-05-18 21:10:49 -07:00
Jared Dillard da7ee68076 support multi-line summaries 2025-05-18 18:17:30 -07:00
Jared Dillard 8ed31f13c4 strip whitespace from summary 2025-05-18 18:11:20 -07:00
Jared Dillard cc38abc8f2 add example links in docs 2025-05-18 17:57:09 -07:00
Jared Dillard bf368670db install pypi version 2025-05-18 17:51:58 -07:00
Jared Dillard a5cbdf15fa rename rtd config file 2025-05-18 17:36:15 -07:00
Jared DillardandGitHub e661ba3da7 Add sphinx docs (#16) 2025-05-18 17:31:00 -07:00
19 changed files with 1584 additions and 250 deletions
+15
View File
@@ -0,0 +1,15 @@
version: 2
build:
os: "ubuntu-20.04"
tools:
python: "3.10"
sphinx:
configuration: docs/source/conf.py
python:
install:
- requirements: docs/requirements.txt
- method: pip
path: .
+35
View File
@@ -1,6 +1,41 @@
Changelog
=========
0.3.2
-----
- Fix image paths to deployed images
`#30 <https://github.com/jdillard/sphinx-llms-txt/pull/30>`_
0.3.1
-----
- Fix issue when ``source_suffix`` equals ``source_link_suffix``
`#29 <https://github.com/jdillard/sphinx-llms-txt/pull/29>`_
0.3.0
-----
- Use first paragraph as default for ``llms_txt_summary``
`#22 <https://github.com/jdillard/sphinx-llms-txt/pull/22>`_
0.2.4
-----
- Support source file suffix detection
`#21 <https://github.com/jdillard/sphinx-llms-txt/pull/21>`_
0.2.3
-----
- Remove ``get_and_resolve_toctree`` method
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
- Simplify ``_sources`` lookup
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
- Add sphinx docs
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
0.2.2
-----
+5 -81
View File
@@ -1,91 +1,15 @@
# Sphinx llms.txt generator
A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText.
A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file.
[![PyPI version](https://img.shields.io/pypi/v/sphinx-llms-txt.svg)](https://pypi.python.org/pypi/sphinx-llms-txt)
[![Conda Version](https://img.shields.io/conda/vn/conda-forge/sphinx-llms-txt.svg)](https://anaconda.org/conda-forge/sphinx-llms-txt)
[![Downloads](https://static.pepy.tech/badge/sphinx-llms-txt/month)](https://pepy.tech/project/sphinx-llms-txt)
[![Parallel Safe](https://img.shields.io/badge/parallel%20safe-true-brightgreen)](#)
## Installation
## Documentation
```bash
pip install sphinx-llms-txt
```
## Usage
1. Add the extension to your Sphinx configuration (`conf.py`):
```python
extensions = [
'sphinx_llms_txt',
]
```
## Configuration Options
### `llms_txt_full_file`
- **Type**: boolean
- **Default**: `'True'`
- **Description**: Whether to write the single output file
### `llms_txt_full_filename`
- **Type**: string
- **Default**: `'llms-full.txt'`
- **Description**: Name of the single output file
### `llms_txt_full_max_size`
- **Type**: integer or `None`
- **Default**: `None` (no limit)
- **Description**: Sets a maximum line count for `llms_txt_full_filename`.
If exceeded, the file is skipped and a warning is shown, but the build still completes.
### `llms_txt_file`
- **Type**: boolean
- **Default**: `True`
- **Description**: Whether to write the summary information file
### `llms_txt_filename`
- **Type**: string
- **Default**: `llms.txt`
- **Description**: Name of the summary information file
### `llms_txt_directives`
- **Type**: list of strings
- **Default**: `[]`
- **Description**: List of custom directive names to process for path resolution.
### `llms_txt_title`
- **Type**: string or `None`
- **Default**: `None`
- **Description**: Overrides the Sphinx project name as the heading in `llms.txt`.
### `llms_txt_summary`
- **Type**: string or `None`
- **Default**: `None`
- **Description**: Optional, but recommended, summary description for `llms.txt`.
### `llms_txt_exclude`
- **Type**: list of strings
- **Default**: `[]`
- **Description**: A list of pages to ignore (e.g., `["page1", "page_with_*"]`).
## Features
- Creates `llms.txt` and `llms-full.txt`
- Automatically add content from `include` directives
- Resolves relative paths in directives like `image` and `figure` to use full paths
- Ability to add list of custom directives with `llms_txt_directives`
- Optionally, prepend a base URL using Sphinx's `html_baseurl`
- Ability to exclude pages
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
## License
+20
View File
@@ -0,0 +1,20 @@
# Minimal makefile for Sphinx documentation
#
# You can set these variables from the command line.
SPHINXOPTS =
SPHINXBUILD = sphinx-build
SPHINXPROJ = SphinxLLMsTxt
SOURCEDIR = source
BUILDDIR = _build
# Put it first so that "make" without argument is like "make help".
help:
@$(SPHINXBUILD) -M help "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
.PHONY: help Makefile
# Catch-all target: route all unknown targets to Sphinx using the new
# "make mode" option. $(O) is meant as a shortcut for $(SPHINXOPTS).
%: Makefile
@$(SPHINXBUILD) -M $@ "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
+6
View File
@@ -0,0 +1,6 @@
furo
esbonio
sphinx-contributors
sphinx
sphinx-llms-txt
sphinxext-opengraph
+170
View File
@@ -0,0 +1,170 @@
Advanced Configuration
======================
This page covers advanced configuration options for the sphinx-llms-txt extension.
.. _customizing_llms_files:
Customizing the LLMs Files
^^^^^^^^^^^^^^^^^^^^^^^^^^
By default, the extension generates two files:
1. ``llms.txt`` - A summary file in Markdown format
2. ``llms-full.txt`` - A complete documentation file in reStructuredText format
You can customize these files in several ways:
.. _changing_filenames:
Changing Filenames
~~~~~~~~~~~~~~~~~~
You can change the default filenames by setting these values in your ``conf.py``:
.. code-block:: python
llms_txt_filename = "custom-summary.txt"
llms_txt_full_filename = "custom-docs.txt"
.. _disabling_file_generation:
Disabling File Generation
~~~~~~~~~~~~~~~~~~~~~~~~~
If you only want one of the files, you can disable generation of the other:
.. code-block:: python
# Disable summary file
llms_txt_file = False
# Disable full documentation file
llms_txt_full_file = False
.. _custom_summary:
Adding a Custom Summary
~~~~~~~~~~~~~~~~~~~~~~~
The summary file can include a custom description of your project:
.. code-block:: python
llms_txt_summary = """
This documentation explains how to use MyProject to build amazing
applications. The project provides a comprehensive API for handling
data processing and visualization.
"""
.. note:: The summary can span multiple lines and will be properly formatted in the output file.
.. _custom_title:
Custom Title
~~~~~~~~~~~~
By default, the project name from Sphinx is used as the title in ``llms.txt``. You can override this:
.. code-block:: python
llms_txt_title = "My Custom Project Documentation"
.. _handling_large_documentation:
Handling Large Documentation
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
You can set a maximum line count:
.. code-block:: python
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
.. _custom_directive_handling:
Custom Directive Handling
^^^^^^^^^^^^^^^^^^^^^^^^^
.. _path_resolution:
Path Resolution
~~~~~~~~~~~~~~~
The extension resolves paths in the common directives ``[ 'image', 'figure']`` by default.
You can add custom directives to this list:
.. code-block:: python
llms_txt_directives = [
"my-custom-image-directive",
"another-directive-with-paths",
]
This ensures that paths in your custom directives are properly resolved in the generated files.
.. _excluding_content:
Excluding Content
^^^^^^^^^^^^^^^^^
You can exclude specific pages from being included in the generated files:
.. code-block:: python
llms_txt_exclude = [
"search", # Exclude the search page
"genindex", # Exclude the index page
"private_*", # Exclude all pages starting with 'private_'
]
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
.. _using_html_baseurl:
Using HTML Base URL
^^^^^^^^^^^^^^^^^^^
If you want to include absolute URLs for resources in your documentation, you can use Sphinx's built-in ``html_baseurl`` configuration:
.. code-block:: python
html_baseurl = "https://example.com/docs/"
When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files.
.. _integration_examples:
Integration Examples
^^^^^^^^^^^^^^^^^^^^
Complete Configuration Example
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Here's a complete example showing multiple :doc:`configuration-values`:
.. code-block:: python
# File names and generation options
llms_txt_filename = "ai-summary.txt"
llms_txt_full_filename = "ai-full-docs.txt"
llms_txt_full_max_size = 50000
# Content customization
llms_txt_title = "Project Documentation for AI Assistants"
llms_txt_summary = """
This is a comprehensive documentation set for our project.
It includes API references, usage examples, and tutorials.
"""
# Path handling
html_baseurl = "https://docs.example.com/"
llms_txt_directives = ["custom-image", "custom-include"]
# Content filtering
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
+1
View File
@@ -0,0 +1 @@
.. include:: ../../CHANGELOG.rst
+105
View File
@@ -0,0 +1,105 @@
#
# Configuration file for the Sphinx documentation builder.
#
# This file does only contain a selection of the most common options. For a
# full list see the documentation:
# http://www.sphinx-doc.org/en/master/config
# -- Path setup --------------------------------------------------------------
import re
import subprocess
# -- Project information -----------------------------------------------------
project = "sphinx-llms-txt"
copyright = "Jared Dillard"
author = "Jared Dillard"
llms_txt_summary = """
A Sphinx extension that generates a summary llms.txt file,written in Markdown,
and a single combined documentation llms-full.txt file, written in reStructuredText.
"""
# check if the current commit is tagged as a release (vX.Y.Z)
try:
GIT_TAG_OUTPUT = subprocess.check_output(["git", "tag", "--points-at", "HEAD"])
current_tag = GIT_TAG_OUTPUT.decode().strip()
if re.match(r"^v(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)$", current_tag):
version = current_tag
else:
version = "latest"
except (subprocess.CalledProcessError, FileNotFoundError):
version = "latest"
# The full version, including alpha/beta/rc tags
release = ""
# -- General configuration ---------------------------------------------------
# If your documentation needs a minimal Sphinx version, state it here.
#
# needs_sphinx = '1.0'
# Add any Sphinx extension module names here, as strings. They can be
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
# ones.
extensions = [
"sphinx.ext.intersphinx",
"sphinx_contributors",
"sphinx_llms_txt",
]
# The language for content autogenerated by Sphinx. Refer to documentation
# for a list of supported languages.
#
# This is also used if you do content translation via gettext catalogs.
# Usually you set "language" from the command line for these cases.
language = "en"
# List of patterns, relative to source directory, that match files and
# directories to ignore when looking for source files.
# This pattern also affects html_static_path and html_extra_path.
exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
# The name of the Pygments (syntax highlighting) style to use.
pygments_style = "sphinx"
intersphinx_mapping = {
"sphinx": ("https://www.sphinx-doc.org/en/master/", None),
}
# -- Options for HTML output -------------------------------------------------
# The theme to use for HTML and HTML Help pages. See the documentation for
# a list of builtin themes.
#
html_theme = "furo"
# Theme options are theme-specific and customize the look and feel of a theme
# further. For a list of options available for each theme, see the
# documentation.
#
html_theme_options = {
"source_repository": "https://github.com/jdillard/sphinx-llms-txt/",
"source_branch": "main",
"source_directory": "docs/source/",
}
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
# -- Options for HTMLHelp output ---------------------------------------------
# Output file base name for HTML help builder.
htmlhelp_basename = "SphinxLLMsTxtdoc"
def setup(app):
app.add_object_type(
"confval",
"confval",
objname="configuration value",
indextemplate="pair: %s; configuration value",
)
+84
View File
@@ -0,0 +1,84 @@
Project Configuration Values
============================
.. confval:: llms_txt_full_file
- **Type**: boolean
- **Default**: ``True``
- **Description**: Whether to write the single output file.
See :ref:`disabling_file_generation`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_full_filename
- **Type**: string
- **Default**: ``'llms-full.txt'``
- **Description**: Name of the single output file.
See :ref:`changing_filenames`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_full_max_size
- **Type**: integer or ``None``
- **Default**: ``None`` (no limit)
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
If exceeded, the file is skipped and a warning is shown, but the build still completes.
See :ref:`handling_large_documentation`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_file
- **Type**: boolean
- **Default**: ``True``
- **Description**: Whether to write the summary information file.
See :ref:`disabling_file_generation`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_filename
- **Type**: string
- **Default**: ``llms.txt``
- **Description**: Name of the summary information file.
See :ref:`changing_filenames`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_directives
- **Type**: list of strings
- **Default**: ``[]`` (empty list)
- **Description**: List of custom directive names to process for path resolution.
See :ref:`path_resolution`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_title
- **Type**: string or ``None``
- **Default**: ``None``
- **Description**: Overrides the Sphinx project name as the heading in ``llms.txt``.
See :ref:`custom_title`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_summary
- **Type**: string
- **Default**: The first paragraph in the root document, else an empty string
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
See :ref:`custom_summary`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_exclude
- **Type**: list of strings
- **Default**: ``[]``
- **Description**: A list of pages to ignore.
See :ref:`excluding_content`.
.. versionadded:: 0.2.1
+48
View File
@@ -0,0 +1,48 @@
Contributing
============
You will need to set up a development environment to make and test your changes before submitting them.
Local development
-----------------
#. Clone the `sphinx-llms-txt repository`_.
#. Create and activate a virtual environment:
.. code-block:: console
python3 -m venv .venv
source .venv/bin/activate
#. Install development dependencies:
.. code-block:: console
pip install -e ".[dev]"
#. Install pre-commit Git hook scripts:
.. code-block:: console
pre-commit install
Testing changes
---------------
Run ``pytest`` before committing changes.
Current contributors
--------------------
Thanks to all who have contributed!
The people that have improved the code:
.. contributors:: jdillard/sphinx-llms-txt
:avatars:
:limit: 100
:exclude: pre-commit-ci[bot],dependabot[bot]
:order: ASC
.. _sphinx-llms-txt repository: https://github.com/jdillard/sphinx-llms-txt
+50
View File
@@ -0,0 +1,50 @@
Getting Started
===============
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Installation
------------
Directly install via ``pip`` by using:
.. code-block:: bash
pip install sphinx-llms-txt
Usage
-----
Add the extension to your Sphinx configuration (``conf.py``):
.. code-block:: python
extensions = [
'sphinx_llms_txt',
]
Once added, the extension will automatically generate the LLMs.txt files during the build process.
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
How It Works
------------
During the Sphinx build process:
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
+34
View File
@@ -0,0 +1,34 @@
Sphinx llms.txt Generator
=========================
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars|
.. toctree::
:maxdepth: 2
getting-started
advanced-configuration
configuration-values
contributing
changelog
.. _Sphinx: http://sphinx-doc.org/
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
:target: https://pypi.python.org/pypi/sphinx-llms-txt
:alt: Latest PyPi Version
.. |Conda Version| image:: https://img.shields.io/conda/vn/conda-forge/sphinx-llms-txt.svg
:target: https://anaconda.org/conda-forge/sphinx-llms-txt
:alt: Latest Conda Version
.. |Downloads| image:: https://static.pepy.tech/badge/sphinx-llms-txt/month
:target: https://pepy.tech/project/sphinx-llms-txt
:alt: PyPi Downloads per month
.. |Parallel Safe| image:: https://img.shields.io/badge/parallel%20safe-true-brightgreen
:target: #
:alt: Parallel read/write safe
.. |GitHub Stars| image:: https://img.shields.io/github/stars/jdillard/sphinx-llms-txt?style=social
:target: https://github.com/jdillard/sphinx-llms-txt
:alt: GitHub Repository stars
+23 -4
View File
@@ -12,7 +12,7 @@ from .manager import LLMSFullManager
from .processor import DocumentProcessor
from .writer import FileWriter
__version__ = "0.2.2"
__version__ = "0.3.2"
# Export classes needed by tests
__all__ = [
@@ -25,9 +25,14 @@ __all__ = [
# Global manager instance
_manager = LLMSFullManager()
# Store root document first paragraph
_root_first_paragraph = ""
def doctree_resolved(app: Sphinx, doctree, docname: str):
"""Called when a docname has been resolved to a document."""
global _root_first_paragraph
# Extract title from the document
title = None
# findall() returns a generator, convert to list to check if it has elements
@@ -38,6 +43,14 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
if title:
_manager.update_page_title(docname, title)
# Extract first paragraph from root document
if docname == app.config.master_doc:
for node in doctree.traverse(nodes.paragraph):
first_para = node.astext()
if first_para:
_root_first_paragraph = first_para
break
def build_finished(app: Sphinx, exception):
"""Called when the build is finished."""
@@ -47,12 +60,17 @@ def build_finished(app: Sphinx, exception):
_manager.set_master_doc(app.config.master_doc)
_manager.set_app(app)
# Get the summary - use configured value or extracted first paragraph
summary = app.config.llms_txt_summary
if summary is None:
summary = _root_first_paragraph
# Set up configuration
config = {
"llms_txt_file": app.config.llms_txt_file,
"llms_txt_filename": app.config.llms_txt_filename,
"llms_txt_title": app.config.llms_txt_title,
"llms_txt_summary": app.config.llms_txt_summary,
"llms_txt_summary": summary,
"llms_txt_full_file": app.config.llms_txt_full_file,
"llms_txt_full_filename": app.config.llms_txt_full_filename,
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
@@ -91,9 +109,10 @@ def setup(app: Sphinx) -> Dict[str, Any]:
app.connect("doctree-resolved", doctree_resolved)
app.connect("build-finished", build_finished)
# Reset manager for each build
global _manager
# Reset manager and root paragraph for each build
global _manager, _root_first_paragraph
_manager = LLMSFullManager()
_root_first_paragraph = ""
return {
"version": __version__,
+124 -25
View File
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
"""
import fnmatch
from typing import Any, Dict, List
from typing import Any, Dict, List, Tuple
from sphinx.environment import BuildEnvironment
from sphinx.util import logging
@@ -19,6 +19,7 @@ class DocumentCollector:
self.master_doc: str = None
self.env: BuildEnvironment = None
self.config: Dict[str, Any] = {}
self.app = None
def set_master_doc(self, master_doc: str):
"""Set the master document name."""
@@ -37,8 +38,79 @@ class DocumentCollector:
"""Set configuration options."""
self.config = config
def get_page_order(self) -> List[str]:
"""Get the correct page order from the toctree structure."""
def set_app(self, app):
"""Set the Sphinx application reference."""
self.app = app
def _get_source_suffixes(self):
"""Get all valid source file suffixes from Sphinx configuration.
Returns:
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
"""
if not self.app:
return [".rst"] # Default fallback
source_suffix = self.app.config.source_suffix
if isinstance(source_suffix, dict):
return list(source_suffix.keys())
elif isinstance(source_suffix, list):
return source_suffix
else:
return [source_suffix] # String format
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
"""
Determine the source suffix for a given docname by checking which
file exists.
Args:
docname: The document name to check
sources_dir: Path to the _sources directory
Returns:
The source suffix if found, or None if no matching file exists
"""
if not sources_dir or not sources_dir.exists():
return None
# Get the source link suffix from Sphinx config
source_link_suffix = ""
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
source_link_suffix = self.app.config.html_sourcelink_suffix
# Handle empty string case specially
if source_link_suffix == "":
source_link_suffix = "" # Keep it empty
elif not source_link_suffix.startswith("."):
source_link_suffix = "." + source_link_suffix
# Get the source file suffixes from Sphinx config
source_suffixes = self._get_source_suffixes()
# Try to find the source file with any of the valid source suffixes
for src_suffix in source_suffixes:
# Avoid duplicate extensions when source_suffix == source_link_suffix
if src_suffix == source_link_suffix:
candidate_file = sources_dir / f"{docname}{src_suffix}"
else:
candidate_file = (
sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
)
if candidate_file.exists():
return src_suffix
return None
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
"""Get the correct page order from the toctree structure.
Args:
sources_dir: Optional path to _sources directory for suffix detection
Returns:
List of tuples (docname, source_suffix) in toctree order
"""
if not self.env or not self.master_doc:
return []
@@ -52,9 +124,12 @@ class DocumentCollector:
visited.add(docname)
# Add the current document
if docname not in page_order:
page_order.append(docname)
# Add the current document with its suffix
if docname not in [doc for doc, _ in page_order]:
suffix = None
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
# Check for toctree entries in this document
try:
@@ -65,20 +140,33 @@ class DocumentCollector:
):
for child_docname in self.env.toctree_includes[docname]:
collect_from_toctree(child_docname)
else:
# Fallback: try to resolve and parse the toctree
toctree = self.env.get_and_resolve_toctree(docname, None)
if toctree:
from docutils import nodes
for node in list(toctree.findall(nodes.reference)):
if "refuri" in node.attributes:
refuri = node.attributes["refuri"]
if refuri and refuri.endswith(".html"):
child_docname = refuri[:-5] # Remove .html
# Try to use dependencies to find related documents
elif (
hasattr(self.env, "dependencies")
and docname in self.env.dependencies
):
# Extract the dependent documents from the dependencies dict
for child_docname in self.env.dependencies[docname]:
# Only add documents actually in the document set
if (
child_docname != docname
): # Avoid circular references
hasattr(self.env, "all_docs")
and child_docname in self.env.all_docs
):
collect_from_toctree(child_docname)
# Fallback to titles or other available references
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
# Get all document names
all_docnames = list(self.env.all_docs.keys())
# Look for documents that might be related (have similar paths)
current_prefix = "/".join(docname.split("/")[:-1])
if current_prefix:
for child_docname in all_docnames:
# Documents in the same directory might be related
if (
child_docname.startswith(current_prefix)
and child_docname != docname
):
collect_from_toctree(child_docname)
except Exception as e:
logger.debug(f"Could not get toctree for {docname}: {e}")
@@ -88,22 +176,33 @@ class DocumentCollector:
# Add any remaining documents not in the toctree (sorted)
if hasattr(self.env, "all_docs"):
processed_docnames = {doc for doc, _ in page_order}
remaining = sorted(
[doc for doc in self.env.all_docs.keys() if doc not in page_order]
[
doc
for doc in self.env.all_docs.keys()
if doc not in processed_docnames
]
)
page_order.extend(remaining)
for docname in remaining:
suffix = None
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
return page_order
def filter_excluded_pages(self, page_order: List[str]) -> List[str]:
def filter_excluded_pages(
self, page_order: List[Tuple[str, str]]
) -> List[Tuple[str, str]]:
"""Filter out excluded pages from the page order."""
exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns:
return [
page
for page in page_order
(docname, suffix)
for docname, suffix in page_order
if not any(
self._match_exclude_pattern(page, pattern)
self._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns
)
]
+135 -92
View File
@@ -56,6 +56,7 @@ class LLMSFullManager:
def set_app(self, app: Sphinx):
"""Set the Sphinx application reference."""
self.app = app
self.collector.set_app(app)
if self.writer:
self.writer.app = app
@@ -69,23 +70,7 @@ class LLMSFullManager:
self.processor = DocumentProcessor(self.config, srcdir)
self.writer = FileWriter(self.config, outdir, self.app)
# Get the correct page order
page_order = self.collector.get_page_order()
if not page_order:
logger.warning(
"Could not determine page order, skipping llms-full creation"
)
return
# Apply exclusion filter if configured
page_order = self.collector.filter_excluded_pages(page_order)
# Determine output file name and location
output_filename = self.config.get("llms_txt_full_filename")
output_path = Path(outdir) / output_filename
# Find sources directory
# Find sources directory first so we can pass it to get_page_order
sources_dir = None
possible_sources = [
Path(outdir) / "_sources",
@@ -104,14 +89,23 @@ class LLMSFullManager:
)
return
# Collect all available source files
txt_files = {}
for f in sources_dir.glob("**/*.txt"):
logger.debug(f"sphinx-llms-txt: Found source file: {f.stem} at {f}")
txt_files[f.stem] = f
# Get the correct page order with source suffixes
page_order = self.collector.get_page_order(sources_dir)
if not page_order:
logger.warning(
"Could not determine page order, skipping llms-full creation"
)
return
# Apply exclusion filter if configured
page_order = self.collector.filter_excluded_pages(page_order)
# Determine output file name and location
output_filename = self.config.get("llms_txt_full_filename")
output_path = Path(outdir) / output_filename
# Log discovered files and page order
logger.debug(f"sphinx-llms-txt: Found {len(txt_files)} source files")
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
# Log exclusion patterns
@@ -119,33 +113,52 @@ class LLMSFullManager:
if exclude_patterns:
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
# Create a mapping from docnames to actual file names
# Create a mapping from docnames to source files
docname_to_file = {}
# Try exact matches first
for docname in page_order:
# Get the source link suffix from Sphinx config
source_link_suffix = (
self.app.config.html_sourcelink_suffix if self.app else ".txt"
)
# Handle empty string case specially
if source_link_suffix == "":
source_link_suffix = "" # Keep it empty
elif not source_link_suffix.startswith("."):
source_link_suffix = "." + source_link_suffix
# Process each (docname, suffix) in the page order
for docname, src_suffix in page_order:
# Skip excluded pages
if any(
if exclude_patterns and any(
self.collector._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns
):
continue
if docname in txt_files:
docname_to_file[docname] = txt_files[docname]
# Build the source file path directly using the known suffix
if src_suffix:
# Avoid duplicate extensions when source_suffix == source_link_suffix
if src_suffix == source_link_suffix:
source_file = sources_dir / f"{docname}{src_suffix}"
expected_suffix = src_suffix
else:
# Try with .rst extension
if f"{docname}.rst" in txt_files:
docname_to_file[docname] = txt_files[f"{docname}.rst"]
# Try with .txt extension
elif f"{docname}.txt" in txt_files:
docname_to_file[docname] = txt_files[f"{docname}.txt"]
# Try with underscores instead of hyphens
elif docname.replace("-", "_") in txt_files:
docname_to_file[docname] = txt_files[docname.replace("-", "_")]
# Try with hyphens instead of underscores
elif docname.replace("_", "-") in txt_files:
docname_to_file[docname] = txt_files[docname.replace("_", "-")]
source_file = (
sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
)
expected_suffix = f"{src_suffix}{source_link_suffix}"
if source_file.exists():
docname_to_file[docname] = source_file
else:
logger.warning(
f"sphinx-llms-txt: Source file not found for: {docname}."
f"Expected: {docname}{expected_suffix}"
)
else:
logger.warning(
f"sphinx-llms-txt: No source suffix determined for: {docname}"
)
# Generate content
content_parts = []
@@ -156,7 +169,7 @@ class LLMSFullManager:
max_lines = self.config.get("llms_txt_full_max_size")
abort_due_to_max_lines = False
for docname in page_order:
for docname, _ in page_order:
if docname in docname_to_file:
file_path = docname_to_file[docname]
content, line_count = self._read_source_file(file_path, docname)
@@ -190,65 +203,77 @@ class LLMSFullManager:
added_files.add(file_path.stem)
total_line_count += line_count
else:
logger.warning(f"sphinx-llm-txt: Source file not found for: {docname}")
logger.warning(
f"sphinx-llms-txt: Source file not found for: {docname}. Check that"
f" file exists at _sources/{docname}[suffix]{source_link_suffix}"
)
# Add any remaining files (in alphabetical order) if not aborted
# Add any remaining files (in alphabetical order) that aren't in the page order
if not abort_due_to_max_lines:
# Apply the same exclusion filter to remaining files
exclude_patterns = self.config.get("llms_txt_exclude")
# Get all source files in the _sources directory using configured suffixes
source_suffixes = self._get_source_suffixes()
all_source_files = []
for src_suffix in source_suffixes:
# Avoid duplicate extensions when source_suffix == source_link_suffix
if src_suffix == source_link_suffix:
glob_pattern = f"**/*{src_suffix}"
else:
glob_pattern = f"**/*{src_suffix}{source_link_suffix}"
all_source_files.extend(sources_dir.glob(glob_pattern))
# Create a set of files to exclude based on their basename
excluded_files = set()
for pattern in exclude_patterns:
if "*" not in pattern and "?" not in pattern:
# For exact patterns, add variants
excluded_files.add(pattern)
excluded_files.add(f"{pattern}.rst")
excluded_files.add(f"{pattern}.txt")
excluded_files.add(pattern.replace("-", "_"))
excluded_files.add(pattern.replace("_", "-"))
processed_paths = set(file.resolve() for file in docname_to_file.values())
# Filter remaining files
remaining_files = sorted(
[
name
for name in txt_files
if name not in added_files
and name not in excluded_files
and not any(
self.collector._match_exclude_pattern(name, pattern)
for pattern in exclude_patterns
)
# Find files that haven't been processed yet
remaining_source_files = [
f for f in all_source_files if f.resolve() not in processed_paths
]
# Sort the remaining files for consistent ordering
remaining_source_files.sort()
if remaining_source_files:
logger.info(
f"Found {len(remaining_source_files)} additional files not in"
f" toctree"
)
if remaining_files:
logger.info(f"Adding remaining files: {remaining_files}")
for file_stem in remaining_files:
file_path = txt_files[file_stem]
content, line_count = self._read_source_file(file_path, file_stem)
for file_path in remaining_source_files:
# Extract docname from path by removing the source and link suffixes
rel_path = str(file_path.relative_to(sources_dir))
docname = None
# Try each source suffix to find which one this file uses
for src_suffix in source_suffixes:
# Avoid duplicate extensions when suffixes match
if src_suffix == source_link_suffix:
combined_suffix = src_suffix
else:
combined_suffix = f"{src_suffix}{source_link_suffix}"
if rel_path.endswith(combined_suffix):
docname = rel_path[: -len(combined_suffix)] # Remove suffix
break
if docname is None:
continue
# Skip excluded docnames
if exclude_patterns and any(
self.collector._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns
):
logger.debug(f"sphinx-llms-txt: Skipping excluded file: {docname}")
continue
# Read and process the file
content, line_count = self._read_source_file(file_path, docname)
# Check if adding this file would exceed the maximum line count
if max_lines is not None and total_line_count + line_count > max_lines:
break
# Double-check that this file should be included
should_include = True
file_stem = file_path.stem
exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns:
# Check stem against exclusion patterns
if any(
self.collector._match_exclude_pattern(file_stem, pattern)
for pattern in exclude_patterns
):
logger.debug(
"sphinx-llms-txt: Final exclusion check removed remaining"
f" file: {file_stem}"
)
should_include = False
if content and should_include:
if content:
logger.debug(f"sphinx-llms-txt: Adding remaining file: {docname}")
content_parts.append(content)
total_line_count += line_count
@@ -258,7 +283,7 @@ class LLMSFullManager:
max_lines is not None and total_line_count > max_lines
):
logger.warning(
f"sphinx-llm-txt: Max line limit ({max_lines}) exceeded:"
f"sphinx-llms-txt: Max line limit ({max_lines}) exceeded:"
f" {total_line_count} > {max_lines}. "
f"Not creating llms-full.txt file."
)
@@ -325,5 +350,23 @@ class LLMSFullManager:
return content_str, line_count + 1
except Exception as e:
logger.error(f"sphinx-llm-txt: Error reading source file {file_path}: {e}")
logger.error(f"sphinx-llms-txt: Error reading source file {file_path}: {e}")
return "", 0
def _get_source_suffixes(self):
"""Get all valid source file suffixes from Sphinx configuration.
Returns:
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
"""
if not self.app:
return [".rst"] # Default fallback
source_suffix = self.app.config.source_suffix
if isinstance(source_suffix, dict):
return list(source_suffix.keys())
elif isinstance(source_suffix, list):
return source_suffix
else:
return [source_suffix] # String format
+48 -6
View File
@@ -94,8 +94,14 @@ class DocumentProcessor:
if not base_url:
return path
# Ensure base URL ends with slash
if not base_url.endswith("/"):
base_url += "/"
# Remove leading slash from path to avoid double slashes
if path.startswith("/"):
path = path[1:]
return f"{base_url}{path}"
def _is_absolute_or_url(self, path: str) -> bool:
@@ -137,8 +143,41 @@ class DocumentProcessor:
prefix = match.group(1) # The entire directive prefix including whitespace
path = match.group(3).strip() # The path argument
# Only process relative paths, not absolute paths or URLs
if not self._is_absolute_or_url(path):
# Handle URLs and data URIs - leave unchanged
if path.startswith(("http://", "https://", "data:")):
return match.group(0)
# For ALL paths, check if image exists in _images first
# Extract filename from the path
filename = os.path.basename(path)
# Check if image exists in _images directory
# First determine the build directory from source_path
build_dir = None
if "_sources" in str(source_path):
# Extract build directory (parent of _sources)
path_parts = str(source_path).split("_sources/")
if len(path_parts) > 1:
build_dir = path_parts[0].rstrip("/")
# If we can determine the build directory, check if image exists in _images
if build_dir:
images_path = os.path.join(build_dir, "_images", filename)
if os.path.exists(images_path):
# Image exists in _images, use _images path
full_path = f"/_images/{filename}"
# Add base URL if configured
full_path = self._add_base_url(full_path, base_url)
return f"{prefix}{full_path}"
# Image doesn't exist in _images, handle based on path type
# Handle absolute paths (starting with /) - add base URL if configured
if path.startswith("/"):
# Add base URL to absolute paths if configured
full_path = self._add_base_url(path, base_url)
return f"{prefix}{full_path}"
# Handle relative paths with original logic for backward compatibility
# Special case for test files
if is_test:
# Add subdir/ prefix to match test expectations
@@ -168,9 +207,7 @@ class DocumentProcessor:
elif rel_doc_dir:
# Join with the original path to form full path relative
# to srcdir
full_path = os.path.normpath(
os.path.join(rel_doc_dir, path)
)
full_path = os.path.normpath(os.path.join(rel_doc_dir, path))
else:
full_path = path
@@ -180,7 +217,12 @@ class DocumentProcessor:
# Return the updated directive with the full path
return f"{prefix}{full_path}"
# If we couldn't resolve the path or it's already absolute, return unchanged
# Fallback for relative paths - add base URL if configured
else:
full_path = self._add_base_url(path, base_url)
return f"{prefix}{full_path}"
# If we couldn't resolve the path, return unchanged
return match.group(0)
# Replace directive paths in the content
+17 -5
View File
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
"""
from pathlib import Path
from typing import Any, Dict, List
from typing import Any, Dict, List, Tuple, Union
from sphinx.application import Sphinx
from sphinx.util import logging
@@ -42,19 +42,19 @@ class FileWriter:
)
return True
except Exception as e:
logger.error(f"sphinx-llm-txt: Error writing combined sources file: {e}")
logger.error(f"sphinx-llms-txt: Error writing combined sources file: {e}")
return False
def write_verbose_info_to_file(
self,
page_order: List[str],
page_order: Union[List[str], List[Tuple[str, str]]],
page_titles: Dict[str, str],
total_line_count: int = 0,
) -> bool:
"""Write summary information to the llms.txt file.
Args:
page_order: Ordered list of document names
page_order: Ordered list of document names or (docname, suffix) tuples
page_titles: Dictionary mapping docnames to titles
total_line_count: Total number of lines in the combined content
@@ -86,6 +86,13 @@ class FileWriter:
# Add description if available
description = self.config.get("llms_txt_summary", "")
if description:
# Trim leading and trailing whitespace
description = description.strip()
if description:
# Only add blockquote if description is not empty
# Replace newlines with newline + blockquote marker to maintain
# blockquote formatting
description = description.replace("\n", "\n> ")
f.write(f"> {description}\n\n")
f.write("## Docs\n\n")
@@ -95,7 +102,12 @@ class FileWriter:
if not base_url.endswith("/"):
base_url += "/"
for docname in page_order:
for item in page_order:
# Handle both old format (str) and new format (tuple)
if isinstance(item, tuple):
docname, _ = item
else:
docname = item
title = page_titles.get(docname, docname)
f.write(f"- [{title}]({base_url}{docname}.html)\n")
+474
View File
@@ -334,3 +334,477 @@ def test_write_verbose_info_with_baseurl(tmp_path):
assert "- [Home Page](https://example.org/index.html)" in content
assert "- [About Us](https://example.org/about.html)" in content
def test_get_source_suffixes_with_dict():
"""Test _get_source_suffixes method with dict source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with dict source_suffix
class MockApp:
class Config:
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert set(suffixes) == {".rst", ".md", ".txt"}
def test_get_source_suffixes_with_list():
"""Test _get_source_suffixes method with list source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with list source_suffix
class MockApp:
class Config:
source_suffix = [".rst", ".md"]
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst", ".md"]
def test_get_source_suffixes_with_string():
"""Test _get_source_suffixes method with string source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with string source_suffix
class MockApp:
class Config:
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"]
def test_get_source_suffixes_no_app():
"""Test _get_source_suffixes method with no app set."""
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"] # Default fallback
def test_html_sourcelink_suffix_default():
"""Test html_sourcelink_suffix defaults to .txt when no app is set."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with default .txt suffix
test_file = f"{sources_dir}/index.rst.txt"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .txt as the default suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed (check if output file exists)
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_custom():
"""Test html_sourcelink_suffix uses custom value from Sphinx config."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with custom html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = "source"
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with custom .source suffix
test_file = f"{sources_dir}/index.rst.source"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .source as the custom suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_with_dot():
"""Test html_sourcelink_suffix adds dot if missing."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with html_sourcelink_suffix without leading dot
class MockApp:
class Config:
html_sourcelink_suffix = "src" # No leading dot
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with .src suffix (dot should be added automatically)
test_file = f"{sources_dir}/index.rst.src"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it adds the dot and finds the file
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_mixed_source_file_formats():
"""Test handling of mixed source file formats (.rst, .md, .txt)."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with multiple source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create test source files with different formats
files_to_create = [
f"{sources_dir}/page1.rst.txt",
f"{sources_dir}/page2.md.txt",
f"{sources_dir}/page3.txt.txt",
]
for test_file in files_to_create:
with open(test_file, "w") as f:
f.write(f"Content for {os.path.basename(test_file)}")
# Mock env with all documents
class MockEnv:
all_docs = {"page1": None, "page2": None, "page3": None}
titles = {
"page1": type("TitleNode", (), {"astext": lambda: "Page 1"})(),
"page2": type("TitleNode", (), {"astext": lambda: "Page 2"})(),
"page3": type("TitleNode", (), {"astext": lambda: "Page 3"})(),
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("page1")
# Test that all file formats are found and processed
manager.combine_sources(outdir, srcdir)
# Verify the output file was created and contains content from all formats
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# Should contain content from all three files
assert "Content for page1.rst.txt" in content
assert "Content for page2.md.txt" in content
assert "Content for page3.txt.txt" in content
def test_source_suffix_detection_priority():
"""Test source suffix detection tries formats in correct order for docnames."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with ordered source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = [".rst", ".md"] # rst has priority over md
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create both .rst and .md versions of the same document
# Only create files for the specific docname "index"
rst_file = f"{sources_dir}/index.rst.txt"
md_file = f"{sources_dir}/index.md.txt"
with open(rst_file, "w") as f:
f.write("RST content for index")
with open(md_file, "w") as f:
f.write("Markdown content for index")
# Mock env with only the index document
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Index Page"})()
}
toctree_includes = {"index": []}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test the priority behavior
manager.combine_sources(outdir, srcdir)
# Check that output file was created
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# The system should prefer RST over MD for the "index" docname
# But since both files exist and the second phase adds remaining files,
# both will be included. The test verifies that RST appears first
# (indicating it was found first in the priority order)
assert "RST content for index" in content
# Find positions to verify order
rst_pos = content.find("RST content for index")
md_pos = content.find("Markdown content for index")
# RST should come before MD (due to priority in toctree processing)
assert rst_pos < md_pos, "RST content should appear before MD content"
def test_summary_default_uses_first_paragraph():
"""
Test that summary defaults to first paragraph of root document when not configured.
"""
from docutils import nodes
from docutils.frontend import OptionParser
from docutils.parsers.rst import Parser
from docutils.utils import new_document
from sphinx_llms_txt import build_finished, doctree_resolved
# Create a proper document with settings
settings = OptionParser(components=(Parser,)).get_default_values()
doctree = new_document("<rst-doc>", settings)
title = nodes.title(text="Test Title")
paragraph = nodes.paragraph(
text="This is the first paragraph that should be used as summary."
)
doctree.append(title)
doctree.append(paragraph)
# Mock Sphinx app
class MockApp:
class Config:
master_doc = "index"
llms_txt_summary = None # Not configured
llms_txt_file = True
llms_txt_filename = "llms.txt"
llms_txt_title = None
llms_txt_full_file = True
llms_txt_full_filename = "llms-full.txt"
llms_txt_full_max_size = None
llms_txt_directives = []
llms_txt_exclude = []
html_baseurl = ""
config = Config()
outdir = "/tmp/build"
srcdir = "/tmp/source"
class Env:
titles = {
"index": type("TitleNode", (), {"astext": lambda self: "Test Title"})()
}
env = Env()
app = MockApp()
# Reset the global state
import sphinx_llms_txt
sphinx_llms_txt._root_first_paragraph = ""
# Call doctree_resolved to extract the first paragraph
doctree_resolved(app, doctree, "index")
# Verify the first paragraph was extracted
assert (
sphinx_llms_txt._root_first_paragraph
== "This is the first paragraph that should be used as summary."
)
# Mock the manager methods to avoid actual file operations
original_combine_sources = sphinx_llms_txt._manager.combine_sources
sphinx_llms_txt._manager.combine_sources = lambda outdir, srcdir: None
# Call build_finished and verify the summary is set correctly
build_finished(app, None)
# Check that the summary was properly configured
assert (
sphinx_llms_txt._manager.config["llms_txt_summary"]
== "This is the first paragraph that should be used as summary."
)
# Restore original method
sphinx_llms_txt._manager.combine_sources = original_combine_sources
+156 -3
View File
@@ -102,7 +102,7 @@ def test_process_path_directives_with_html_baseurl(tmp_path):
def test_process_path_directives_absolute_urls(tmp_path):
"""Test that absolute URLs are not modified."""
"""Test that absolute URLs are not modified but absolute paths get base URL."""
# Create a processor
config = {
"llms_txt_directives": [],
@@ -127,10 +127,17 @@ def test_process_path_directives_absolute_urls(tmp_path):
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives (should remain unchanged)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
assert processed_content == source_content
# Expected: URLs and data URIs unchanged, absolute paths get base URL
expected_content = (
".. image:: https://othersite.com/images/test.png\n"
".. image:: https://example.com/docs/absolute/path/image.png\n"
".. image:: data:image/png;base64,iVBORw0KG...\n"
)
assert processed_content == expected_content
def test_process_path_directives_custom_directives(tmp_path):
@@ -251,3 +258,149 @@ def test_process_content_end_to_end(tmp_path):
)
assert processed_content == expected_content
def test_process_path_directives_images_directory(tmp_path):
"""Test that _images directory paths are handled correctly."""
# Create a processor with base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "https://example.com/docs",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create _sources directory to mimic Sphinx output
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Create a source file with various _images directory paths
source_content = (
"Some content.\n"
".. image:: _images/test.png\n" # Relative _images should become /_images
".. image:: /_images/absolute.png\n" # Absolute _images should get base URL
".. figure:: _images/figure.png\n" # Test with figure directive too
" :alt: A test figure\n"
".. image:: images/normal.png\n" # Normal relative path should be unchanged
)
# Create source file in sources directory to simulate Sphinx build output
source_file = sources_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: _images paths should be converted and get base URL
expected_content = (
"Some content.\n"
".. image:: https://example.com/docs/_images/test.png\n"
".. image:: https://example.com/docs/_images/absolute.png\n"
".. figure:: https://example.com/docs/_images/figure.png\n"
" :alt: A test figure\n"
".. image:: https://example.com/docs/images/normal.png\n"
)
assert processed_content == expected_content
def test_process_path_directives_images_directory_no_baseurl(tmp_path):
"""
Test that _images directory paths work correctly without base URL.
Only converts when image exists.
"""
# Create a processor without base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create _sources directory to mimic Sphinx output
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Create _images directory and one test image
images_dir = build_dir / "_images"
images_dir.mkdir()
(images_dir / "test.png").write_text("fake image content")
# Note: absolute.png is not created, so it won't be converted
# Create a source file with _images directory paths
source_content = (
".. image:: _images/test.png\n" # Should become /_images (image exists)
".. image:: /_images/absolute.png\n" # Should stay unchanged (absolute path)
)
# Create source file in sources directory to simulate Sphinx build output
source_file = sources_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: only test.png gets converted because it exists in _images
expected_content = (
".. image:: /_images/test.png\n" # Converted because image exists
".. image:: /_images/absolute.png\n" # Absolute path unchanged
)
assert processed_content == expected_content
def test_process_path_directives_all_absolute_paths_get_baseurl(tmp_path):
"""Test that all absolute paths (starting with /) get base URL prepended."""
# Create a processor with base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "https://mysite.com/docs/",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create a source file with various absolute paths
source_content = (
".. image:: /static/images/logo.png\n"
".. figure:: /assets/diagrams/flow.svg\n"
".. image:: /media/photos/team.jpg\n"
" :alt: Team photo\n"
".. image:: relative/path.png\n" # This should still get normal processing
)
# Create source file
source_file = src_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: All absolute paths get base URL prepended
expected_content = (
".. image:: https://mysite.com/docs/static/images/logo.png\n"
".. figure:: https://mysite.com/docs/assets/diagrams/flow.svg\n"
".. image:: https://mysite.com/docs/media/photos/team.jpg\n"
" :alt: Team photo\n"
".. image:: https://mysite.com/docs/relative/path.png\n"
)
assert processed_content == expected_content