cherry2md.py - converter Cherrytree notes to markdown [🌳📄] -----> [⚙️] -----> [📝]
Table of Contents
Conversion of 'rich text' to markdown is a not so unambiguously job. For instance, what to do when some text was marked up with monospace font and also marked up in italic style? The choice that I made was to prefer monospace markup over all other styles that are present; monospace text will be converted as 'inline code' or as a 'code block' (depending on the presence of newlines in the text). Reason is that I personally choosed many times for monospace (just press CTRL-m in Cherrytree) over a codebox, I just found that easier to do (regretting my laziness afterwards).
I have made almost 500 notes in Cherrytree and it is was surprising to see that many notes did not match the format rules I came up with. So this converter does not deal with all possible situations and I accept to modify some markdown notes manually after the conversion.
Cherrytree exports all images as encoded BASE64 in 'png' file format. Images will be stored together in one 'images' directory whereby each image will be suffixed with an unique number.'''
The program reads the XML-database file exported by Cherrytree and converts all the nodes to files and directories analogous to the hierarchy within the Cherrytree database.
This tool is written in Python (version 3.12) using standard libraries only.
Take a look at the command-line options and the TOML file in the sections that follow (see also my recommendations in the README section on GitHub).
These steps show how to install and run cherry2md in a clean and isolated environment. It is recommended to keep your development tools inside a dedicated directory (e.g., ~/myprojects) and to use a Python virtual environment.
Choose or create a directory for your personal projects:
mkdir ~/myprojects # Linux/MAc
cd ~/myprojectsgit clone https://github.com/rikroos/cherry2markdown.git
cd cherry2markdownCreating a virtual environment prevents conflicts with other Python tools or system packages.
python3 -m venv venv
source venv/bin/activate # Linux/macOS
# or on Windows:
# venv\Scripts\activateUpgrade pip (recommended):
pip install --upgrade pipYou can now view the available command-line options:
python app/cherry2md.py --helpCreate a sub-directory to store the XM-files exported from Cherrytree, then move the XML-file(s) into this new subdir:
mkdir xml-filesTo convert a Cherrytree XML database:
python app/cherry2md.py <path-to-xml-file>Verbose mode examples:
python app/cherry2md.py <path-to-xml-file> -v # verbose
python app/cherry2md.py <path-to-xml-file> -vv # very verboseStarting at the root only the tags <node> are selected.
Within a <node> tag other <node> tags are recursively processed.
Other tags inside a <node> tag that will be processed are:
| Tag | Descriptions |
|---|---|
<rich_text> |
regular tags with text |
<codebox> |
programming code |
<encode_png> |
base64 encoded images but also other text- objects like PDF-files |
<table> |
text in columns like a regular table |
<link> |
links to web-sites or other resources |
Only Cherrytree notes with some content are legible to be written as markdown file; thus empty notes are silently discarded. This can be however an issue for cases where the existance of the note-node it self is functioning as some form of documentation. In that case it is advised to drop a character into the empty Cherrytree note e.g. a "!" or a "." before exporting the XML-file.
In case an empty node is referencing one or multiple sub-nodes then these sub-nodes will be placed in a newly created directory that wil be named after the title of that parent-node; non-alphabetical chars are cleansed in that name.
The tag <rich_text> has attributes that formats the text in a Cherrytree note.
These attributes has to be translated into some arbitrary markdown formats.
-
Attribute "scale" with values "h1" till "h6" : Exported as "#" till "######".
-
Attribute "family" with value "monospace" : Textfragment is exported surrounded by ```.
-
How superscript and subscript are handled: markdown (as I know) does not support these out of the box. I choose to translate these to matching UTF-8 characters.
Only tags <node> that contain one or more tags <rich_text> are written to a markdown file. This markdown file is named after the name of the node; spaces are replaced with underscores.
In case a tag <node> contains other tags <node> a new directory will be created with the name of the parent node (without the ".md" extension). But this is only done if at least one child node contains text fragment(s) which will result in a new note file for that child.
If the parent <node> it self also contains text fragment(s) then a new note file is created as a sibling of the created directory that will hold the child notes of this parent <node>.
So, a branch in the tree of Cherrytree notes without any text are silently discarded from output (no empty note files will be generated)! See also paragraph 1 which commented on this case.
Cherrytree has a few buttons to transform selected lines into lists with the desired list-item symbol. This will result in plain text in the text fragment which will be exported as is (UTF-characters).
• item 1 • item 2 • item 3
☐ item 1 ☐ item 2 ☐ item 3
- item 1
- item 2
- item 3
In Markdown line-ends are processed as if not existing. So lines are concatenated together. Only an empty line is respected; however, the following empty lines are then ignored. But I wanted to migrate these additional empty lines too.
A simple trick is to append 2 spaces to each line-end. This way you don't have to worry about the line-ends.
In the config directory you will find a TOML-file with a few settings that dictates how to deal with line-ends. Default 2 spaces are appended to each line found in the XML-file to mimic the layout of the orignal note layout.
After finishing the conversion a list with some counts will be printed but only when commandline parameter -v or -vv is used . A note about the count of 'created directories': only created directories containing a real markdown-note are counted for. Reason is that directories are created with the option 'parents=True'.
When running in DRY-RUN mode, afterwards a message will be printed to warn you nothing has been created.
The converter can be customized using a TOML file (for content-related settings) and command-line parameters (for runtime behavior).
Locatie: app/config/cherry2md.toml
| Key | Description |
|---|---|
| skip_notes | ignore Cherrytree notes with the specified IDs |
| images_link | root path for the image location (a dedicated subdirectory in Obsidian) |
| pdf_link | root path for PDF files (a dedicated subdirectory in Obsidian) |
| others_link | root path for other file types (a dedicated subdirectory in Obsidian) |
| horizontal_rule | markup |
| italic | markup |
| heavy | markup |
| strikethrough | markup |
| background_color | markup |
| superscript | markup |
| subscript | markup |
| missing_linklabel.max_chars | default: 99 — maximum number of characters for auto-generated link text |
| missing_linklabel.shortened_text | default: "...(more)" — text appended when a label is shortened |
| fmt_timestamp_created | "%Y-%m-%d %H:%M:%S" |
| fmt_timestamp_exported | "%Y-%m-%d %H:%M:%S %z" |
| append_each_line | default: two spaces, which forces a line break |
| replace_empty_line_with | default: "" (empty), or e.g. "\\" or <br> |
| note.title | default: enabled — optional prefix placed before the original note text |
| note.closing | default: enabled — optional postfix placed after the original note (export metadata) |
Obsidian has a feature that automatically assigns each note a # chapter title based on the Markdown filename.
If this is not desired and the user disables that feature, the converter can be configured to generate a title instead.
In the TOML file the prefix text is currently set to only a divider line ---. You can embed the variable {title} to display the generated title.
The converter always requires at least one argument: the path to the XML file that should be converted.
$ python cherry2md.py <path-to-xml-file>Additional options can be provided, for example to display extra information about the results:
$ python cherry2md.py <path-to-xml-file> -vOr to display even more verbose output:
$ python cherry2md.py <path-to-xml-file> -vvThe available runtime command-line arguments can be listed with:
$ python cherry2md.py --help| argument | descriptions |
|---|---|
xml_filepath |
the path and filename of the Cherrytree XML file to process; this can be an absolute or a relative path |
| argument | descriptions |
|---|---|
-h, --help |
Show this help message and exit. |
-a, --absolute_path |
References inside markdown notes: use absolute path instead of relative paths in references to other files like images. Default a relative path is generated because that makes it possible to move the generated notes-tree to a other place in the filesystem afterwards. In case you like to browse your notes-tree in a sandwiched environment, for example with 'jail', you can opt in for an absolute path. |
-c, --clean_env |
Clean up the output environment: purge the directory 'markdown' from the 'data_dir' directory. You will be prompted for confirmation [y/n]."-> Be carefull not to run this option after having made changes manually to your newly created markdown notes! |
-d, --data_dir DATA_DIR |
Directory for containing the output, will be created if not existing already. default: './data/_ch2md_' .For security reasons, the data directory is always followed by a system generated directory called ch2md |
-x, --attachment_index ATTACHMENT_INDEX |
Starting number for indexing attachment files; default is 1. When doing multiple runs, do not forget to use this option with the correct index-no to start the next batch. |
-q, --quiet |
No output information on screen but redirect all info to file 'stdout.log' |
-u, --unique_id UNIQUE_ID |
Convert only a specific Cherrytree-note with the Cherrytree-note unique_id (a XML-attribute) |
-v, --verbosity |
Increase output verbosity: -v : this shows used directory paths and a table with counts -vv : this shows additionally filepaths of created notes and directories |
-y, --dryrun |
fake run: disable the creation of any file on disk |
The Author, November 2025