Compare commits
14
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
935a964c7e | ||
|
|
51f6c71de3 | ||
|
|
70defd3996 | ||
|
|
9ae05c6c13 | ||
|
|
5581979cac | ||
|
|
f70f1a26ec | ||
|
|
ed50138ae4 | ||
|
|
8f4d2c07c6 | ||
|
|
da7ee68076 | ||
|
|
8ed31f13c4 | ||
|
|
cc38abc8f2 | ||
|
|
bf368670db | ||
|
|
a5cbdf15fa | ||
|
|
e661ba3da7 |
@@ -0,0 +1,15 @@
|
|||||||
|
version: 2
|
||||||
|
|
||||||
|
build:
|
||||||
|
os: "ubuntu-20.04"
|
||||||
|
tools:
|
||||||
|
python: "3.10"
|
||||||
|
|
||||||
|
sphinx:
|
||||||
|
configuration: docs/source/conf.py
|
||||||
|
|
||||||
|
python:
|
||||||
|
install:
|
||||||
|
- requirements: docs/requirements.txt
|
||||||
|
- method: pip
|
||||||
|
path: .
|
||||||
@@ -1,6 +1,16 @@
|
|||||||
Changelog
|
Changelog
|
||||||
=========
|
=========
|
||||||
|
|
||||||
|
0.2.3
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Remove ``get_and_resolve_toctree`` method
|
||||||
|
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
|
||||||
|
- Simplify ``_sources`` lookup
|
||||||
|
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
|
||||||
|
- Add sphinx docs
|
||||||
|
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
|
||||||
|
|
||||||
0.2.2
|
0.2.2
|
||||||
-----
|
-----
|
||||||
|
|
||||||
|
|||||||
@@ -1,91 +1,13 @@
|
|||||||
# Sphinx llms.txt generator
|
# Sphinx llms.txt generator
|
||||||
|
|
||||||
A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText.
|
A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file.
|
||||||
|
|
||||||
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
||||||
[](https://pepy.tech/project/sphinx-llms-txt)
|
[](https://pepy.tech/project/sphinx-llms-txt)
|
||||||
|
|
||||||
## Installation
|
## Documentation
|
||||||
|
|
||||||
```bash
|
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
||||||
pip install sphinx-llms-txt
|
|
||||||
```
|
|
||||||
|
|
||||||
## Usage
|
|
||||||
|
|
||||||
1. Add the extension to your Sphinx configuration (`conf.py`):
|
|
||||||
|
|
||||||
```python
|
|
||||||
extensions = [
|
|
||||||
'sphinx_llms_txt',
|
|
||||||
]
|
|
||||||
```
|
|
||||||
|
|
||||||
## Configuration Options
|
|
||||||
|
|
||||||
### `llms_txt_full_file`
|
|
||||||
|
|
||||||
- **Type**: boolean
|
|
||||||
- **Default**: `'True'`
|
|
||||||
- **Description**: Whether to write the single output file
|
|
||||||
|
|
||||||
### `llms_txt_full_filename`
|
|
||||||
|
|
||||||
- **Type**: string
|
|
||||||
- **Default**: `'llms-full.txt'`
|
|
||||||
- **Description**: Name of the single output file
|
|
||||||
|
|
||||||
### `llms_txt_full_max_size`
|
|
||||||
|
|
||||||
- **Type**: integer or `None`
|
|
||||||
- **Default**: `None` (no limit)
|
|
||||||
- **Description**: Sets a maximum line count for `llms_txt_full_filename`.
|
|
||||||
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
|
||||||
|
|
||||||
### `llms_txt_file`
|
|
||||||
|
|
||||||
- **Type**: boolean
|
|
||||||
- **Default**: `True`
|
|
||||||
- **Description**: Whether to write the summary information file
|
|
||||||
|
|
||||||
### `llms_txt_filename`
|
|
||||||
|
|
||||||
- **Type**: string
|
|
||||||
- **Default**: `llms.txt`
|
|
||||||
- **Description**: Name of the summary information file
|
|
||||||
|
|
||||||
### `llms_txt_directives`
|
|
||||||
|
|
||||||
- **Type**: list of strings
|
|
||||||
- **Default**: `[]`
|
|
||||||
- **Description**: List of custom directive names to process for path resolution.
|
|
||||||
|
|
||||||
### `llms_txt_title`
|
|
||||||
|
|
||||||
- **Type**: string or `None`
|
|
||||||
- **Default**: `None`
|
|
||||||
- **Description**: Overrides the Sphinx project name as the heading in `llms.txt`.
|
|
||||||
|
|
||||||
### `llms_txt_summary`
|
|
||||||
|
|
||||||
- **Type**: string or `None`
|
|
||||||
- **Default**: `None`
|
|
||||||
- **Description**: Optional, but recommended, summary description for `llms.txt`.
|
|
||||||
|
|
||||||
### `llms_txt_exclude`
|
|
||||||
|
|
||||||
- **Type**: list of strings
|
|
||||||
- **Default**: `[]`
|
|
||||||
- **Description**: A list of pages to ignore (e.g., `["page1", "page_with_*"]`).
|
|
||||||
|
|
||||||
## Features
|
|
||||||
|
|
||||||
- Creates `llms.txt` and `llms-full.txt`
|
|
||||||
- Automatically add content from `include` directives
|
|
||||||
- Resolves relative paths in directives like `image` and `figure` to use full paths
|
|
||||||
- Ability to add list of custom directives with `llms_txt_directives`
|
|
||||||
- Optionally, prepend a base URL using Sphinx's `html_baseurl`
|
|
||||||
- Ability to exclude pages
|
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
# Minimal makefile for Sphinx documentation
|
||||||
|
#
|
||||||
|
|
||||||
|
# You can set these variables from the command line.
|
||||||
|
SPHINXOPTS =
|
||||||
|
SPHINXBUILD = sphinx-build
|
||||||
|
SPHINXPROJ = SphinxLLMsTxt
|
||||||
|
SOURCEDIR = source
|
||||||
|
BUILDDIR = _build
|
||||||
|
|
||||||
|
# Put it first so that "make" without argument is like "make help".
|
||||||
|
help:
|
||||||
|
@$(SPHINXBUILD) -M help "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
|
||||||
|
|
||||||
|
.PHONY: help Makefile
|
||||||
|
|
||||||
|
# Catch-all target: route all unknown targets to Sphinx using the new
|
||||||
|
# "make mode" option. $(O) is meant as a shortcut for $(SPHINXOPTS).
|
||||||
|
%: Makefile
|
||||||
|
@$(SPHINXBUILD) -M $@ "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
furo
|
||||||
|
esbonio
|
||||||
|
sphinx-contributors
|
||||||
|
sphinx
|
||||||
|
sphinx-llms-txt
|
||||||
|
sphinxext-opengraph
|
||||||
@@ -0,0 +1,170 @@
|
|||||||
|
Advanced Configuration
|
||||||
|
======================
|
||||||
|
|
||||||
|
This page covers advanced configuration options for the sphinx-llms-txt extension.
|
||||||
|
|
||||||
|
.. _customizing_llms_files:
|
||||||
|
|
||||||
|
Customizing the LLMs Files
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
By default, the extension generates two files:
|
||||||
|
|
||||||
|
1. ``llms.txt`` - A summary file in Markdown format
|
||||||
|
2. ``llms-full.txt`` - A complete documentation file in reStructuredText format
|
||||||
|
|
||||||
|
You can customize these files in several ways:
|
||||||
|
|
||||||
|
.. _changing_filenames:
|
||||||
|
|
||||||
|
Changing Filenames
|
||||||
|
~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
You can change the default filenames by setting these values in your ``conf.py``:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_filename = "custom-summary.txt"
|
||||||
|
llms_txt_full_filename = "custom-docs.txt"
|
||||||
|
|
||||||
|
.. _disabling_file_generation:
|
||||||
|
|
||||||
|
Disabling File Generation
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
If you only want one of the files, you can disable generation of the other:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# Disable summary file
|
||||||
|
llms_txt_file = False
|
||||||
|
|
||||||
|
# Disable full documentation file
|
||||||
|
llms_txt_full_file = False
|
||||||
|
|
||||||
|
.. _custom_summary:
|
||||||
|
|
||||||
|
Adding a Custom Summary
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
The summary file can include a custom description of your project:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_summary = """
|
||||||
|
This documentation explains how to use MyProject to build amazing
|
||||||
|
applications. The project provides a comprehensive API for handling
|
||||||
|
data processing and visualization.
|
||||||
|
"""
|
||||||
|
|
||||||
|
.. note:: The summary can span multiple lines and will be properly formatted in the output file.
|
||||||
|
|
||||||
|
.. _custom_title:
|
||||||
|
|
||||||
|
Custom Title
|
||||||
|
~~~~~~~~~~~~
|
||||||
|
|
||||||
|
By default, the project name from Sphinx is used as the title in ``llms.txt``. You can override this:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_title = "My Custom Project Documentation"
|
||||||
|
|
||||||
|
.. _handling_large_documentation:
|
||||||
|
|
||||||
|
Handling Large Documentation
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
||||||
|
You can set a maximum line count:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
||||||
|
|
||||||
|
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
|
||||||
|
|
||||||
|
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
|
||||||
|
|
||||||
|
.. _custom_directive_handling:
|
||||||
|
|
||||||
|
Custom Directive Handling
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
.. _path_resolution:
|
||||||
|
|
||||||
|
Path Resolution
|
||||||
|
~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
The extension resolves paths in the common directives ``[ 'image', 'figure']`` by default.
|
||||||
|
You can add custom directives to this list:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_directives = [
|
||||||
|
"my-custom-image-directive",
|
||||||
|
"another-directive-with-paths",
|
||||||
|
]
|
||||||
|
|
||||||
|
This ensures that paths in your custom directives are properly resolved in the generated files.
|
||||||
|
|
||||||
|
.. _excluding_content:
|
||||||
|
|
||||||
|
Excluding Content
|
||||||
|
^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
You can exclude specific pages from being included in the generated files:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_exclude = [
|
||||||
|
"search", # Exclude the search page
|
||||||
|
"genindex", # Exclude the index page
|
||||||
|
"private_*", # Exclude all pages starting with 'private_'
|
||||||
|
]
|
||||||
|
|
||||||
|
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
||||||
|
|
||||||
|
.. _using_html_baseurl:
|
||||||
|
|
||||||
|
Using HTML Base URL
|
||||||
|
^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
If you want to include absolute URLs for resources in your documentation, you can use Sphinx's built-in ``html_baseurl`` configuration:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
html_baseurl = "https://example.com/docs/"
|
||||||
|
|
||||||
|
When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files.
|
||||||
|
|
||||||
|
.. _integration_examples:
|
||||||
|
|
||||||
|
Integration Examples
|
||||||
|
^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
Complete Configuration Example
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
Here's a complete example showing multiple :ref:`configuration-values`:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# File names and generation options
|
||||||
|
llms_txt_filename = "ai-summary.txt"
|
||||||
|
llms_txt_full_filename = "ai-full-docs.txt"
|
||||||
|
llms_txt_full_max_size = 50000
|
||||||
|
|
||||||
|
# Content customization
|
||||||
|
llms_txt_title = "Project Documentation for AI Assistants"
|
||||||
|
llms_txt_summary = """
|
||||||
|
This is a comprehensive documentation set for our project.
|
||||||
|
It includes API references, usage examples, and tutorials.
|
||||||
|
"""
|
||||||
|
|
||||||
|
# Path handling
|
||||||
|
html_baseurl = "https://docs.example.com/"
|
||||||
|
llms_txt_directives = ["custom-image", "custom-include"]
|
||||||
|
|
||||||
|
# Content filtering
|
||||||
|
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
.. include:: ../../CHANGELOG.rst
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
#
|
||||||
|
# Configuration file for the Sphinx documentation builder.
|
||||||
|
#
|
||||||
|
# This file does only contain a selection of the most common options. For a
|
||||||
|
# full list see the documentation:
|
||||||
|
# http://www.sphinx-doc.org/en/master/config
|
||||||
|
|
||||||
|
# -- Path setup --------------------------------------------------------------
|
||||||
|
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
|
||||||
|
# -- Project information -----------------------------------------------------
|
||||||
|
|
||||||
|
project = "sphinx-llms-txt"
|
||||||
|
copyright = "Jared Dillard"
|
||||||
|
author = "Jared Dillard"
|
||||||
|
llms_txt_summary = """
|
||||||
|
A Sphinx extension that generates a summary llms.txt file,written in Markdown,
|
||||||
|
and a single combined documentation llms-full.txt file, written in reStructuredText.
|
||||||
|
"""
|
||||||
|
|
||||||
|
# check if the current commit is tagged as a release (vX.Y.Z)
|
||||||
|
try:
|
||||||
|
GIT_TAG_OUTPUT = subprocess.check_output(["git", "tag", "--points-at", "HEAD"])
|
||||||
|
current_tag = GIT_TAG_OUTPUT.decode().strip()
|
||||||
|
if re.match(r"^v(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)$", current_tag):
|
||||||
|
version = current_tag
|
||||||
|
else:
|
||||||
|
version = "latest"
|
||||||
|
except (subprocess.CalledProcessError, FileNotFoundError):
|
||||||
|
version = "latest"
|
||||||
|
|
||||||
|
# The full version, including alpha/beta/rc tags
|
||||||
|
release = ""
|
||||||
|
|
||||||
|
|
||||||
|
# -- General configuration ---------------------------------------------------
|
||||||
|
|
||||||
|
# If your documentation needs a minimal Sphinx version, state it here.
|
||||||
|
#
|
||||||
|
# needs_sphinx = '1.0'
|
||||||
|
|
||||||
|
# Add any Sphinx extension module names here, as strings. They can be
|
||||||
|
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
|
||||||
|
# ones.
|
||||||
|
extensions = [
|
||||||
|
"sphinx.ext.intersphinx",
|
||||||
|
"sphinx_contributors",
|
||||||
|
"sphinx_llms_txt",
|
||||||
|
]
|
||||||
|
|
||||||
|
# The language for content autogenerated by Sphinx. Refer to documentation
|
||||||
|
# for a list of supported languages.
|
||||||
|
#
|
||||||
|
# This is also used if you do content translation via gettext catalogs.
|
||||||
|
# Usually you set "language" from the command line for these cases.
|
||||||
|
language = "en"
|
||||||
|
|
||||||
|
# List of patterns, relative to source directory, that match files and
|
||||||
|
# directories to ignore when looking for source files.
|
||||||
|
# This pattern also affects html_static_path and html_extra_path.
|
||||||
|
exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
|
||||||
|
|
||||||
|
# The name of the Pygments (syntax highlighting) style to use.
|
||||||
|
pygments_style = "sphinx"
|
||||||
|
|
||||||
|
intersphinx_mapping = {
|
||||||
|
"sphinx": ("https://www.sphinx-doc.org/en/master/", None),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# -- Options for HTML output -------------------------------------------------
|
||||||
|
|
||||||
|
# The theme to use for HTML and HTML Help pages. See the documentation for
|
||||||
|
# a list of builtin themes.
|
||||||
|
#
|
||||||
|
html_theme = "furo"
|
||||||
|
|
||||||
|
# Theme options are theme-specific and customize the look and feel of a theme
|
||||||
|
# further. For a list of options available for each theme, see the
|
||||||
|
# documentation.
|
||||||
|
#
|
||||||
|
html_theme_options = {}
|
||||||
|
|
||||||
|
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
|
||||||
|
|
||||||
|
|
||||||
|
# -- Options for HTMLHelp output ---------------------------------------------
|
||||||
|
|
||||||
|
# Output file base name for HTML help builder.
|
||||||
|
htmlhelp_basename = "SphinxLLMsTxtdoc"
|
||||||
|
|
||||||
|
|
||||||
|
def setup(app):
|
||||||
|
app.add_object_type(
|
||||||
|
"confval",
|
||||||
|
"confval",
|
||||||
|
objname="configuration value",
|
||||||
|
indextemplate="pair: %s; configuration value",
|
||||||
|
)
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
Project Configuration Values
|
||||||
|
============================
|
||||||
|
|
||||||
|
.. confval:: llms_txt_full_file
|
||||||
|
|
||||||
|
- **Type**: boolean
|
||||||
|
- **Default**: ``True``
|
||||||
|
- **Description**: Whether to write the single output file.
|
||||||
|
See :ref:`disabling_file_generation`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_full_filename
|
||||||
|
|
||||||
|
- **Type**: string
|
||||||
|
- **Default**: ``'llms-full.txt'``
|
||||||
|
- **Description**: Name of the single output file.
|
||||||
|
See :ref:`changing_filenames`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_full_max_size
|
||||||
|
|
||||||
|
- **Type**: integer or ``None``
|
||||||
|
- **Default**: ``None`` (no limit)
|
||||||
|
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
||||||
|
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
||||||
|
See :ref:`handling_large_documentation`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_file
|
||||||
|
|
||||||
|
- **Type**: boolean
|
||||||
|
- **Default**: ``True``
|
||||||
|
- **Description**: Whether to write the summary information file.
|
||||||
|
See :ref:`disabling_file_generation`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_filename
|
||||||
|
|
||||||
|
- **Type**: string
|
||||||
|
- **Default**: ``llms.txt``
|
||||||
|
- **Description**: Name of the summary information file.
|
||||||
|
See :ref:`changing_filenames`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_directives
|
||||||
|
|
||||||
|
- **Type**: list of strings
|
||||||
|
- **Default**: ``[]`` (empty list)
|
||||||
|
- **Description**: List of custom directive names to process for path resolution.
|
||||||
|
See :ref:`path_resolution`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_title
|
||||||
|
|
||||||
|
- **Type**: string or ``None``
|
||||||
|
- **Default**: ``None``
|
||||||
|
- **Description**: Overrides the Sphinx project name as the heading in ``llms.txt``.
|
||||||
|
See :ref:`custom_title`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_summary
|
||||||
|
|
||||||
|
- **Type**: string or ``None``
|
||||||
|
- **Default**: ``None``
|
||||||
|
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
|
||||||
|
See :ref:`custom_summary`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_exclude
|
||||||
|
|
||||||
|
- **Type**: list of strings
|
||||||
|
- **Default**: ``[]``
|
||||||
|
- **Description**: A list of pages to ignore.
|
||||||
|
See :ref:`excluding_content`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.2.1
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
Contributing
|
||||||
|
============
|
||||||
|
|
||||||
|
You will need to set up a development environment to make and test your changes before submitting them.
|
||||||
|
|
||||||
|
Local development
|
||||||
|
-----------------
|
||||||
|
|
||||||
|
#. Clone the `sphinx-llms-txt repository`_.
|
||||||
|
|
||||||
|
#. Create and activate a virtual environment:
|
||||||
|
|
||||||
|
.. code-block:: console
|
||||||
|
|
||||||
|
python3 -m venv .venv
|
||||||
|
source .venv/bin/activate
|
||||||
|
|
||||||
|
#. Install development dependencies:
|
||||||
|
|
||||||
|
.. code-block:: console
|
||||||
|
|
||||||
|
pip install -e ".[dev]"
|
||||||
|
|
||||||
|
#. Install pre-commit Git hook scripts:
|
||||||
|
|
||||||
|
.. code-block:: console
|
||||||
|
|
||||||
|
pre-commit install
|
||||||
|
|
||||||
|
Testing changes
|
||||||
|
---------------
|
||||||
|
|
||||||
|
Run ``pytest`` before committing changes.
|
||||||
|
|
||||||
|
Current contributors
|
||||||
|
--------------------
|
||||||
|
|
||||||
|
Thanks to all who have contributed!
|
||||||
|
The people that have improved the code:
|
||||||
|
|
||||||
|
.. contributors:: jdillard/sphinx-llms-txt
|
||||||
|
:avatars:
|
||||||
|
:limit: 100
|
||||||
|
:exclude: pre-commit-ci[bot],dependabot[bot]
|
||||||
|
:order: ASC
|
||||||
|
|
||||||
|
|
||||||
|
.. _sphinx-llms-txt repository: https://github.com/jdillard/sphinx-llms-txt
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
Getting Started
|
||||||
|
===============
|
||||||
|
|
||||||
|
Demo
|
||||||
|
----
|
||||||
|
|
||||||
|
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
||||||
|
|
||||||
|
Installation
|
||||||
|
------------
|
||||||
|
|
||||||
|
Directly install via ``pip`` by using:
|
||||||
|
|
||||||
|
.. code-block:: bash
|
||||||
|
|
||||||
|
pip install sphinx-llms-txt
|
||||||
|
|
||||||
|
Usage
|
||||||
|
-----
|
||||||
|
|
||||||
|
Add the extension to your Sphinx configuration (``conf.py``):
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
extensions = [
|
||||||
|
'sphinx_llms_txt',
|
||||||
|
]
|
||||||
|
|
||||||
|
Once added, the extension will automatically generate the LLMs.txt files during the build process.
|
||||||
|
|
||||||
|
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
||||||
|
|
||||||
|
How It Works
|
||||||
|
-----------
|
||||||
|
|
||||||
|
During the Sphinx build process:
|
||||||
|
|
||||||
|
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
|
||||||
|
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
||||||
|
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
||||||
|
4. **Output Generation**: Creates two optional files:
|
||||||
|
|
||||||
|
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
||||||
|
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
||||||
|
|
||||||
|
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
|
||||||
|
|
||||||
|
|
||||||
|
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
||||||
|
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
Sphinx llms.txt Generator
|
||||||
|
=========================
|
||||||
|
|
||||||
|
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|
||||||
|
|
||||||
|
|PyPI version|
|
||||||
|
|
||||||
|
.. toctree::
|
||||||
|
:maxdepth: 2
|
||||||
|
|
||||||
|
getting-started
|
||||||
|
advanced-configuration
|
||||||
|
configuration-values
|
||||||
|
contributing
|
||||||
|
changelog
|
||||||
|
|
||||||
|
|
||||||
|
.. _Sphinx: http://sphinx-doc.org/
|
||||||
|
|
||||||
|
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
||||||
|
:target: https://pypi.python.org/pypi/sphinx-llms-txt
|
||||||
|
:alt: Latest PyPi Version
|
||||||
@@ -12,7 +12,7 @@ from .manager import LLMSFullManager
|
|||||||
from .processor import DocumentProcessor
|
from .processor import DocumentProcessor
|
||||||
from .writer import FileWriter
|
from .writer import FileWriter
|
||||||
|
|
||||||
__version__ = "0.2.2"
|
__version__ = "0.2.3"
|
||||||
|
|
||||||
# Export classes needed by tests
|
# Export classes needed by tests
|
||||||
__all__ = [
|
__all__ = [
|
||||||
@@ -58,7 +58,6 @@ def build_finished(app: Sphinx, exception):
|
|||||||
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
||||||
"llms_txt_directives": app.config.llms_txt_directives,
|
"llms_txt_directives": app.config.llms_txt_directives,
|
||||||
"llms_txt_exclude": app.config.llms_txt_exclude,
|
"llms_txt_exclude": app.config.llms_txt_exclude,
|
||||||
"llms_txt_rm_directives": app.config.llms_txt_rm_directives,
|
|
||||||
"html_baseurl": getattr(app.config, "html_baseurl", ""),
|
"html_baseurl": getattr(app.config, "html_baseurl", ""),
|
||||||
}
|
}
|
||||||
_manager.set_config(config)
|
_manager.set_config(config)
|
||||||
@@ -87,7 +86,6 @@ def setup(app: Sphinx) -> Dict[str, Any]:
|
|||||||
app.add_config_value("llms_txt_title", None, "env")
|
app.add_config_value("llms_txt_title", None, "env")
|
||||||
app.add_config_value("llms_txt_summary", None, "env")
|
app.add_config_value("llms_txt_summary", None, "env")
|
||||||
app.add_config_value("llms_txt_exclude", [], "env")
|
app.add_config_value("llms_txt_exclude", [], "env")
|
||||||
app.add_config_value("llms_txt_rm_directives", False, "env")
|
|
||||||
|
|
||||||
# Connect to Sphinx events
|
# Connect to Sphinx events
|
||||||
app.connect("doctree-resolved", doctree_resolved)
|
app.connect("doctree-resolved", doctree_resolved)
|
||||||
|
|||||||
@@ -65,21 +65,34 @@ class DocumentCollector:
|
|||||||
):
|
):
|
||||||
for child_docname in self.env.toctree_includes[docname]:
|
for child_docname in self.env.toctree_includes[docname]:
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(child_docname)
|
||||||
else:
|
# Try to use dependencies to find related documents
|
||||||
# Fallback: try to resolve and parse the toctree
|
elif (
|
||||||
toctree = self.env.get_and_resolve_toctree(docname, None)
|
hasattr(self.env, "dependencies")
|
||||||
if toctree:
|
and docname in self.env.dependencies
|
||||||
from docutils import nodes
|
):
|
||||||
|
# Extract the dependent documents from the dependencies dict
|
||||||
|
for child_docname in self.env.dependencies[docname]:
|
||||||
|
# Only add documents actually in the document set
|
||||||
|
if (
|
||||||
|
hasattr(self.env, "all_docs")
|
||||||
|
and child_docname in self.env.all_docs
|
||||||
|
):
|
||||||
|
collect_from_toctree(child_docname)
|
||||||
|
# Fallback to titles or other available references
|
||||||
|
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
|
||||||
|
# Get all document names
|
||||||
|
all_docnames = list(self.env.all_docs.keys())
|
||||||
|
|
||||||
for node in list(toctree.findall(nodes.reference)):
|
# Look for documents that might be related (have similar paths)
|
||||||
if "refuri" in node.attributes:
|
current_prefix = "/".join(docname.split("/")[:-1])
|
||||||
refuri = node.attributes["refuri"]
|
if current_prefix:
|
||||||
if refuri and refuri.endswith(".html"):
|
for child_docname in all_docnames:
|
||||||
child_docname = refuri[:-5] # Remove .html
|
# Documents in the same directory might be related
|
||||||
if (
|
if (
|
||||||
child_docname != docname
|
child_docname.startswith(current_prefix)
|
||||||
): # Avoid circular references
|
and child_docname != docname
|
||||||
collect_from_toctree(child_docname)
|
):
|
||||||
|
collect_from_toctree(child_docname)
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.debug(f"Could not get toctree for {docname}: {e}")
|
logger.debug(f"Could not get toctree for {docname}: {e}")
|
||||||
|
|
||||||
|
|||||||
+53
-74
@@ -104,14 +104,7 @@ class LLMSFullManager:
|
|||||||
)
|
)
|
||||||
return
|
return
|
||||||
|
|
||||||
# Collect all available source files
|
|
||||||
txt_files = {}
|
|
||||||
for f in sources_dir.glob("**/*.txt"):
|
|
||||||
logger.debug(f"sphinx-llms-txt: Found source file: {f.stem} at {f}")
|
|
||||||
txt_files[f.stem] = f
|
|
||||||
|
|
||||||
# Log discovered files and page order
|
# Log discovered files and page order
|
||||||
logger.debug(f"sphinx-llms-txt: Found {len(txt_files)} source files")
|
|
||||||
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
||||||
|
|
||||||
# Log exclusion patterns
|
# Log exclusion patterns
|
||||||
@@ -119,33 +112,28 @@ class LLMSFullManager:
|
|||||||
if exclude_patterns:
|
if exclude_patterns:
|
||||||
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
|
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
|
||||||
|
|
||||||
# Create a mapping from docnames to actual file names
|
# Create a mapping from docnames to source files
|
||||||
docname_to_file = {}
|
docname_to_file = {}
|
||||||
|
|
||||||
# Try exact matches first
|
# Process each docname in the page order
|
||||||
for docname in page_order:
|
for docname in page_order:
|
||||||
# Skip excluded pages
|
# Skip excluded pages
|
||||||
if any(
|
if exclude_patterns and any(
|
||||||
self.collector._match_exclude_pattern(docname, pattern)
|
self.collector._match_exclude_pattern(docname, pattern)
|
||||||
for pattern in exclude_patterns
|
for pattern in exclude_patterns
|
||||||
):
|
):
|
||||||
continue
|
continue
|
||||||
|
|
||||||
if docname in txt_files:
|
# Construct expected source file path directly from docname
|
||||||
docname_to_file[docname] = txt_files[docname]
|
source_file = sources_dir / f"{docname}.rst.txt"
|
||||||
|
|
||||||
|
if source_file.exists():
|
||||||
|
docname_to_file[docname] = source_file
|
||||||
else:
|
else:
|
||||||
# Try with .rst extension
|
logger.warning(
|
||||||
if f"{docname}.rst" in txt_files:
|
f"sphinx-llm-txt: Source file not found for: {docname}. Expected"
|
||||||
docname_to_file[docname] = txt_files[f"{docname}.rst"]
|
f" at {source_file}"
|
||||||
# Try with .txt extension
|
)
|
||||||
elif f"{docname}.txt" in txt_files:
|
|
||||||
docname_to_file[docname] = txt_files[f"{docname}.txt"]
|
|
||||||
# Try with underscores instead of hyphens
|
|
||||||
elif docname.replace("-", "_") in txt_files:
|
|
||||||
docname_to_file[docname] = txt_files[docname.replace("-", "_")]
|
|
||||||
# Try with hyphens instead of underscores
|
|
||||||
elif docname.replace("_", "-") in txt_files:
|
|
||||||
docname_to_file[docname] = txt_files[docname.replace("_", "-")]
|
|
||||||
|
|
||||||
# Generate content
|
# Generate content
|
||||||
content_parts = []
|
content_parts = []
|
||||||
@@ -190,65 +178,56 @@ class LLMSFullManager:
|
|||||||
added_files.add(file_path.stem)
|
added_files.add(file_path.stem)
|
||||||
total_line_count += line_count
|
total_line_count += line_count
|
||||||
else:
|
else:
|
||||||
logger.warning(f"sphinx-llm-txt: Source file not found for: {docname}")
|
logger.warning(
|
||||||
|
f"sphinx-llm-txt: Source file not found for: {docname}. Check that"
|
||||||
|
f" the file exists at _sources/{docname}.rst.txt"
|
||||||
|
)
|
||||||
|
|
||||||
# Add any remaining files (in alphabetical order) if not aborted
|
# Add any remaining files (in alphabetical order) that aren't in the page order
|
||||||
if not abort_due_to_max_lines:
|
if not abort_due_to_max_lines:
|
||||||
# Apply the same exclusion filter to remaining files
|
# Get all .rst.txt files in the _sources directory
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
all_source_files = list(sources_dir.glob("**/*.rst.txt"))
|
||||||
|
processed_paths = set(file.resolve() for file in docname_to_file.values())
|
||||||
|
|
||||||
# Create a set of files to exclude based on their basename
|
# Find files that haven't been processed yet
|
||||||
excluded_files = set()
|
remaining_source_files = [
|
||||||
for pattern in exclude_patterns:
|
f for f in all_source_files if f.resolve() not in processed_paths
|
||||||
if "*" not in pattern and "?" not in pattern:
|
]
|
||||||
# For exact patterns, add variants
|
|
||||||
excluded_files.add(pattern)
|
|
||||||
excluded_files.add(f"{pattern}.rst")
|
|
||||||
excluded_files.add(f"{pattern}.txt")
|
|
||||||
excluded_files.add(pattern.replace("-", "_"))
|
|
||||||
excluded_files.add(pattern.replace("_", "-"))
|
|
||||||
|
|
||||||
# Filter remaining files
|
# Sort the remaining files for consistent ordering
|
||||||
remaining_files = sorted(
|
remaining_source_files.sort()
|
||||||
[
|
|
||||||
name
|
if remaining_source_files:
|
||||||
for name in txt_files
|
logger.info(
|
||||||
if name not in added_files
|
f"Found {len(remaining_source_files)} additional files not in"
|
||||||
and name not in excluded_files
|
f" toctree"
|
||||||
and not any(
|
)
|
||||||
self.collector._match_exclude_pattern(name, pattern)
|
|
||||||
for pattern in exclude_patterns
|
for file_path in remaining_source_files:
|
||||||
)
|
# Extract docname from path by removing the .rst.txt extension
|
||||||
]
|
rel_path = str(file_path.relative_to(sources_dir))
|
||||||
)
|
if rel_path.endswith(".rst.txt"):
|
||||||
if remaining_files:
|
docname = rel_path[:-8] # Remove .rst.txt extension
|
||||||
logger.info(f"Adding remaining files: {remaining_files}")
|
else:
|
||||||
for file_stem in remaining_files:
|
continue
|
||||||
file_path = txt_files[file_stem]
|
|
||||||
content, line_count = self._read_source_file(file_path, file_stem)
|
# Skip excluded docnames
|
||||||
|
if exclude_patterns and any(
|
||||||
|
self.collector._match_exclude_pattern(docname, pattern)
|
||||||
|
for pattern in exclude_patterns
|
||||||
|
):
|
||||||
|
logger.debug(f"sphinx-llms-txt: Skipping excluded file: {docname}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
# Read and process the file
|
||||||
|
content, line_count = self._read_source_file(file_path, docname)
|
||||||
|
|
||||||
# Check if adding this file would exceed the maximum line count
|
# Check if adding this file would exceed the maximum line count
|
||||||
if max_lines is not None and total_line_count + line_count > max_lines:
|
if max_lines is not None and total_line_count + line_count > max_lines:
|
||||||
break
|
break
|
||||||
|
|
||||||
# Double-check that this file should be included
|
if content:
|
||||||
should_include = True
|
logger.debug(f"sphinx-llms-txt: Adding remaining file: {docname}")
|
||||||
file_stem = file_path.stem
|
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
|
||||||
|
|
||||||
if exclude_patterns:
|
|
||||||
# Check stem against exclusion patterns
|
|
||||||
if any(
|
|
||||||
self.collector._match_exclude_pattern(file_stem, pattern)
|
|
||||||
for pattern in exclude_patterns
|
|
||||||
):
|
|
||||||
logger.debug(
|
|
||||||
"sphinx-llms-txt: Final exclusion check removed remaining"
|
|
||||||
f" file: {file_stem}"
|
|
||||||
)
|
|
||||||
should_include = False
|
|
||||||
|
|
||||||
if content and should_include:
|
|
||||||
content_parts.append(content)
|
content_parts.append(content)
|
||||||
total_line_count += line_count
|
total_line_count += line_count
|
||||||
|
|
||||||
|
|||||||
@@ -50,33 +50,8 @@ class DocumentProcessor:
|
|||||||
# Then process path directives (image, figure, etc.)
|
# Then process path directives (image, figure, etc.)
|
||||||
content = self._process_path_directives(content, source_path)
|
content = self._process_path_directives(content, source_path)
|
||||||
|
|
||||||
# Remove directives if configured to do so
|
|
||||||
if self.config.get("llms_txt_rm_directives", False):
|
|
||||||
content = self._remove_directives(content)
|
|
||||||
|
|
||||||
return content
|
return content
|
||||||
|
|
||||||
def _remove_directives(self, content: str) -> str:
|
|
||||||
"""Remove directives from content.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
content: The source content from which to remove directives
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Content with all directives removed
|
|
||||||
"""
|
|
||||||
# Match any directive pattern (starting with .. followed by ::)
|
|
||||||
directive_pattern = re.compile(r'^\s*\.\.\s+[\w\-]+::.*?$(?:\n\s+.*?$)*',
|
|
||||||
re.MULTILINE | re.DOTALL)
|
|
||||||
|
|
||||||
# Replace all directives with an empty string
|
|
||||||
processed_content = directive_pattern.sub('', content)
|
|
||||||
|
|
||||||
# Clean up any consecutive blank lines that might result from directive removal
|
|
||||||
processed_content = re.sub(r'\n{3,}', '\n\n', processed_content)
|
|
||||||
|
|
||||||
return processed_content
|
|
||||||
|
|
||||||
def _extract_relative_document_path(
|
def _extract_relative_document_path(
|
||||||
self, source_path: Path
|
self, source_path: Path
|
||||||
) -> Tuple[Optional[str], Optional[str], Optional[List[str]]]:
|
) -> Tuple[Optional[str], Optional[str], Optional[List[str]]]:
|
||||||
|
|||||||
@@ -86,6 +86,11 @@ class FileWriter:
|
|||||||
# Add description if available
|
# Add description if available
|
||||||
description = self.config.get("llms_txt_summary", "")
|
description = self.config.get("llms_txt_summary", "")
|
||||||
if description:
|
if description:
|
||||||
|
# Trim leading and trailing whitespace
|
||||||
|
description = description.strip()
|
||||||
|
# Replace newlines with newline + blockquote marker to maintain
|
||||||
|
# blockquote formatting
|
||||||
|
description = description.replace("\n", "\n> ")
|
||||||
f.write(f"> {description}\n\n")
|
f.write(f"> {description}\n\n")
|
||||||
|
|
||||||
f.write("## Docs\n\n")
|
f.write("## Docs\n\n")
|
||||||
|
|||||||
@@ -334,44 +334,3 @@ def test_write_verbose_info_with_baseurl(tmp_path):
|
|||||||
|
|
||||||
assert "- [Home Page](https://example.org/index.html)" in content
|
assert "- [Home Page](https://example.org/index.html)" in content
|
||||||
assert "- [About Us](https://example.org/about.html)" in content
|
assert "- [About Us](https://example.org/about.html)" in content
|
||||||
|
|
||||||
|
|
||||||
def test_remove_directives():
|
|
||||||
"""Test removing directives from content."""
|
|
||||||
# Create a processor with remove_directives enabled
|
|
||||||
config = {"llms_txt_rm_directives": True}
|
|
||||||
processor = DocumentProcessor(config)
|
|
||||||
|
|
||||||
# Test content with various directives
|
|
||||||
content = """This is a test document.
|
|
||||||
|
|
||||||
.. image:: /path/to/image.jpg
|
|
||||||
:alt: An example image
|
|
||||||
:width: 100%
|
|
||||||
|
|
||||||
This is a paragraph after the image.
|
|
||||||
|
|
||||||
.. note::
|
|
||||||
This is a note.
|
|
||||||
|
|
||||||
.. code-block:: python
|
|
||||||
|
|
||||||
def hello_world():
|
|
||||||
print("Hello, world!")
|
|
||||||
|
|
||||||
Final paragraph."""
|
|
||||||
|
|
||||||
processed_content = processor._remove_directives(content)
|
|
||||||
|
|
||||||
# Check that directives are removed
|
|
||||||
assert ".. image::" not in processed_content
|
|
||||||
assert ".. note::" not in processed_content
|
|
||||||
assert ".. code-block::" not in processed_content
|
|
||||||
|
|
||||||
# Check that regular content is preserved
|
|
||||||
assert "This is a test document." in processed_content
|
|
||||||
assert "This is a paragraph after the image." in processed_content
|
|
||||||
assert "Final paragraph." in processed_content
|
|
||||||
|
|
||||||
# Check that there are no excessive blank lines
|
|
||||||
assert "\n\n\n" not in processed_content
|
|
||||||
|
|||||||
Reference in New Issue
Block a user