Hello World from Org
First published: 25 May 2025
I describe yet-another-simple-publishing-setup for a static website using Emacs's Org mode. While there are plenty of resources on the topic available already, I wanted to document my setup here for self-reference. Maybe someone will take away an idea or two.
skywhi@dreamland:~/org/blog$ echo "Hello World Wild Web!" > src/posts/hello.org
skywhi@dreamland:~/org/blog$ make update && make publish
skywhi@dreamland:~/org/blog$ # make blogging great again!
Contents:
Rationale#
I figured the time had come to host some of my modest content online for sharing, fun, persistence, easy access and all the other usual suspects. Going into this adventure there were only a few requirements that I wanted my website to meet:
- Content should be written with a
cursedtext editor and Org mode. - No need for any back-end or front-end logic: serving a static website without any fancy HTML/CSS/JS shenanigans will be more than fine.
- The website should be standalone so that it can be viewed offline as well.
- Convenient hosting: administrating a full-blown webserver for such a small project would be unnecessary hassle.
Let me explain. I have been using Emacs and Org mode for several years now. Its markup language really fits my needs and I have integrated my Org files with the agenda and other nifty features Emacs and Org have to offer. The ability to hack their behavior in elisp is also a great plus to me as it immensely increases the flexibility of this setup.
When it comes to this website, I want to be able to turn whatever I am currently working on into publishable content without having to spend extra time on re-formatting my writing from the ground up. It was thus pretty obvious that I would give a try to the built-in org-publish package for Emacs as it is exactly what I am looking for: automated and configurable exporting of Org projects to HTML files, as plain as possible. There exist quite a few other projects which turn a collection of Org files into a website, such as Hugo or org-static-blog. But I figured I would like to retain maximal control over the exporting process and keep external dependencies as low as possible so I stuck to org-publish.
Using git versioning gives me the opportunity to track and manage changes over time. Since we are dealing with a static website with exclusively public data, GitHub Pages will handle the hosting for me, under my own domain name. Whenever I want to work on a new item, I can pull the remote github repository, create a local branch and start editing. When the content is ready for publishing, the branch is merged into main, pushed to the remote github repository and voilĂ ! More on the deployment part further down.
Project layout#
I ended up with something like the following directory structure for this project. The github repository containing all the up-to-date source files can be found here.
- ~/org/blog/
- .github/
- workflows/
- static.yml
- static/
- css/
- fonts/
- img/
- js/
- mathjax/
- other/
- src/
- index.org
- about.org
- links.org
- now.org
- 404.org
- notes/
- index.org
- a-note-i-keep-poking-at.org
- posts/
- files/
- hello.org
- super-dupa-post.org
- templates/
- header.html
- footer.html
- post.org
- www/
- CNAME
- LICENSE
- Makefile
- publish-website.el
- README.org
src/ contains the Org files that will be published to HTML. These files declare as few exporting options as possible so that they can easily be used in other contexts or projects. Global exporting options will be defined in publish-website.el and page-specific options provided via templates/ or, in rare cases, directly in the files themselves. The static/ directory hosts assets such as stylesheets, fonts, images, etc. that will be exported "as is". www/ contains the complete exported website and – this one matters – it is committed along with the sources, for reasons explained in the deployment section.
The logic behind publishing is implemented in publish-website.el which is called by a Makefile for convenience. It defines all the directories, files and exporting options that will be fed into org-publish. I could have written my publish-website.el in literate programming style and use it as source for this blog post, but that would have been a little bit too meta, even for me :)
Finally, README.org is just a symbolic link to src/about.org. Why? Because we can!
To the first page!#
Source files#
This is for example the beginning of this post. It does not contain much besides
a title, a date, a template and a table of content. There is also a custom
preview block, but more on that a little bit later.
#+TITLE: Hello World from Org
#+DATE: [2025-05-25 Sun]
#+begin_preview
I describe /yet-another-simple-publishing-setup/ for a static website using
Emacs's Org mode. While there are plenty of resources on the topic available
already, I wanted to document my setup here for self-reference. Maybe someone
will take away an idea or two.
#+end_preview
#+begin_src sh
skywhi@dreamland:~/org/blog$ echo "Hello World Wild Web!" > src/posts/hello.org
skywhi@dreamland:~/org/blog$ make update && make publish
skywhi@dreamland:~/org/blog$ # make blogging great again!
#+end_src
{{{toc(Contents:,6)}}}
* Rationale
:PROPERTIES:
:CUSTOM_ID: rationale
:END:
I figured the time had come to host some of my modest content online for
sharing, fun, persistence, easy access and all the other usual suspects. Going
into this adventure there were only a few requirements that I wanted my website
to meet: [...]
Appending a CUSTOM_ID property (C-c C-x p) to headings makes it possible to
link to them in the HTML. This example link points to the present section. Headings
carrying one also render a clickable # so that the link can be grabbed from
the page itself, see further down.
The toc macro above expands to a bold caption followed by a #+TOC keyword so
that every page declares its table of content the same way. It is defined along
with the other macros further down.
Templates#
When several Org files share the same rendering and exporting options, it can
make sense to regroup them in an Org template file and use the #+SETUPFILE or
#+INCLUDE directive in the relevant source files.
For example, the template under templates/post.org will extract the date
directive of the source file and add custom HTML to the post.
@@html:<span class="post-date">@@
/First published: {{{date(%d %b %Y)}}}/
@@html:</span>@@
CSS styling#
The two files that are used are style.css and htmlize.css. The first hosts all
the base styles for the website while the second prettifies source code blocks
by redefining various .org-* CSS classes. Since I am aiming for a small and
flexible website I first wrote my own CSS by hand – and I suck at it tbh. After
some consideration, and out of respect for the general public's good taste, I
eventually handed the design over to the AI overlords. Anyone who would rather
not go through the same painful wonderful experience should look for
pre-existing stylesheets and nice exporting options instead; some good
org-compatible ones are available out there.
Building and publishing#
publish-website.el#
It was hyped for long enough, here is the publish-website.el file. This file defines common settings and build instructions for the project. We feed them into org-publish which will take care of producing the corresponding website.
The file has grown to a bit under 600 lines, but most of it is plain
configuration and publishing rules; only a handful of functions really need a
word of explanation, and those are the ones reproduced along the way. I started
with the two below. They generate a "preview" of each post, visible in the
list of all posts, and prepend a formatted timestamp of when the post was first
added – a custom macro takes care of the formatting. The content of the
"preview" is whatever sits between the custom #+begin_preview and
#+end_preview tags.
;; Preview block
(defun skw-blog/get-preview (file)
"Extract the content between #+begin_preview and #+end_preview blocks
in 'file'. The block tags have to be on their own lines, preferably
before and after paragraphs. Return an empty string when 'file' has
no such block."
(with-temp-buffer
(message file)
(insert-file-contents file)
(goto-char (point-min))
(let* ((beg (and (re-search-forward "^#\\+begin_preview$" nil t)
(+ 1 (point))))
(end (and beg
(re-search-forward "^#\\+end_preview$" nil t)
(match-beginning 0))))
(if end
(replace-regexp-in-string "\n" " " (buffer-substring beg end))
""))))
;; Format list of blog post for the sitemap / index
(defun skw-blog/org-format-blog-post (entry style project)
"Format 'entry' in org-publish 'project' sitemap to include a timestamp."
(let ((entry-title (org-publish-find-title entry project)))
(if (= (length entry-title) 0)
(format "*%s*" entry)
(format "{{{timestamp(%s)}}}: [[file:%s][%s]]"
(format-time-string "%Y-%m-%d" (org-publish-find-date entry project))
entry
entry-title))))
;; Same but add the content between the "preview" tags
(defun skw-blog/org-format-blog-post-with-preview (entry style project)
"Format 'entry' in org-publish 'project' sitemap to include a timestamp
and preview ('begin/end_preview' tag)."
(let ((entry-title (org-publish-find-title entry project))
(preview (skw-blog/get-preview (concat (skw-blog/get-root-directory) "src/posts/" entry)))) ;; dirty
(if (= (length entry-title) 0)
(format "*%s*" entry)
(format "{{{timestamp(%s)}}}: [[file:%s][%s]]\n
%s"
(format-time-string "%Y-%m-%d" (org-publish-find-date entry project))
entry
entry-title
preview))))
;; Exporting macros
(setq org-export-global-macros
'(("timestamp" . "@@html:<span class=\"timestamp\">$1</span>@@")
("toc" . "*$1*\n#+TOC: headlines $2 local")))
The preview block itself is optional. When a file does not have one, both
searches are allowed to fail and skw-blog/get-preview simply hands back an
empty string.
The site header lives in templates/header.html, whose path is held by
skw-blog/header-file. It is read once into skw-blog/header and injected as
the HTML preamble. Rather than serving the same markup everywhere, a small
function marks the navigation entry matching the page being exported with
aria-current so that the current section can be styled (and announced by
screen readers).
(defun skw-blog/get-file-content (file)
"Return the content of 'file' as a string"
(with-temp-buffer
(insert-file-contents file)
(buffer-string)))
(defvar skw-blog/header
(skw-blog/get-file-content skw-blog/header-file))
(defun skw-blog/nav-href-for (input-file)
"Return the header nav href matching 'input-file', or nil when none does."
(let ((f (or input-file "")))
(cond ((string-match-p "/posts/" f) "/posts/")
((string-match-p "/notes/" f) "/notes/")
((string-match-p "links\\.org\\'" f) "/links.html")
((string-match-p "now\\.org\\'" f) "/now.html")
((string-match-p "about\\.org\\'" f) "/about.html"))))
(defun skw-blog/html-preamble (info)
"Return the site header, marking the nav entry for the page being exported."
(let ((href (skw-blog/nav-href-for (plist-get info :input-file))))
(if (null href)
skw-blog/header
(replace-regexp-in-string
(concat "<a href=\"" (regexp-quote href) "\"")
(concat "<a href=\"" href "\" aria-current=\"page\"")
skw-blog/header t t))))
Wide tables used to drag the whole page sideways on narrow screens. An export
filter wraps every exported table in its own scrollable container, so only the
table scrolls. The tabindex makes that container reachable with the keyboard.
(defun skw-blog/wrap-table-in-scroll-container (table backend info)
"Wrap an exported HTML 'table' in a horizontally scrollable container."
(if (org-export-derived-backend-p backend 'html)
(format "<div class=\"table-scroll\" tabindex=\"0\">\n%s</div>\n" table)
table))
(add-to-list 'org-export-filter-table-functions
#'skw-blog/wrap-table-in-scroll-container)
Three publishing projects turn src/ into pages. website-src does the
exporting: it walks src/ recursively, runs every Org file through
org-html-publish-to-html and adds the shared header, footer and <head> along
the way. The other two only generate an Org file to be used as an index for blog
posts, with and without previews, hence their skw-blog/publish-nothing.
(defun skw-blog/publish-nothing (_plist _filename _dir)
"Publish nothing. Used by projects that only exist to generate a sitemap."
nil)
(defun skw-blog/format-latest-posts (title list)
"Format the 'skw-blog/latest-posts-count' first entries of 'list' as a
sitemap titled 'title'."
(let ((entries (seq-take (cdr list) skw-blog/latest-posts-count)))
(concat "#+TITLE: " title "\n\n"
(org-list-to-org (cons (car list) entries)))))
("website-src"
:auto-sitemap t
:base-directory ,skw-blog/srcdir
:base-extension "org"
:exclude ,(regexp-opt '("rss.org" "index-no-preview.org"))
:html-head ,(concat skw-blog/main-css skw-blog/favicon skw-blog/rss-link)
:html-postamble ,skw-blog/footer
:html-preamble skw-blog/html-preamble
:publishing-directory ,skw-blog/outdir
:publishing-function org-html-publish-to-html
:recursive t
:sitemap-function skw-blog/format-main-sitemap
:sitemap-title ,skw-blog/main-sitemap-title)
;; Bare list of recent posts, included by the homepage. Never a page of its own.
("website-posts-index"
:auto-sitemap t
:base-directory ,(concat skw-blog/srcdir "posts")
:base-extension "org"
:exclude ,(regexp-opt '("rss.org" "index.org" "index-no-preview.org"))
:publishing-directory ,(concat skw-blog/outdir "posts")
:publishing-function skw-blog/publish-nothing
:sitemap-filename "index-no-preview.org"
:sitemap-format-entry skw-blog/org-format-blog-post
:sitemap-function skw-blog/format-latest-posts
:sitemap-sort-files anti-chronologically
:sitemap-title "Latest posts")
;; Full post list with previews. website-src is what turns it into a page.
("website-posts-index-preview"
:auto-sitemap t
:base-directory ,(concat skw-blog/srcdir "posts")
:base-extension "org"
:exclude ,(regexp-opt '("rss.org" "index.org" "index-no-preview.org"))
:publishing-directory ,(concat skw-blog/outdir "posts")
:publishing-function skw-blog/publish-nothing
:sitemap-filename "index.org"
:sitemap-format-entry skw-blog/org-format-blog-post-with-preview
:sitemap-sort-files anti-chronologically
:sitemap-title "Posts")
The site-wide sitemap gets a :sitemap-function of its own for one small
touch-up: it drops the 404 page from the listing.
(defun skw-blog/format-main-sitemap (title list)
"Format the site-wide sitemap, leaving the 404 page out of 'list'.
That page answers every unknown path, so listing it as a destination of its
own only invites visitors and crawlers to walk into it."
(replace-regexp-in-string
"^[ \t]*-[ \t]*\\[\\[file:404\\.org\\]\\[[^]]*\\]\\][ \t]*\n" ""
(org-publish-sitemap-default title list)))
The generated src/posts/index-no-preview.org is inserted into the main src/index.org file like this:
Org files are not the only items that have to end up in www/. Stylesheets,
fonts, images and the MathJax install are copied over verbatim by two more
projects, both relying on the built-in org-publish-attachment function rather
than on an exporter. website-files picks up the assets that sit next to the
sources under src/, while website-static takes care of everything under
static/.
;; Attachment files
("website-files"
:base-directory ,skw-blog/srcdir
:base-extension "css\\|txt\\|jpg\\|gif\\|png"
:publishing-directory ,skw-blog/outdir
:publishing-function org-publish-attachment
:recursive t)
;; Static files
("website-static"
:base-directory ,(concat skw-blog/rootdir "static")
:base-extension ".*"
:exclude "\\.org\\'"
:publishing-directory ,(concat skw-blog/outdir "static")
:publishing-function org-publish-attachment
:recursive t)
Order matters in the :components list: the two index projects have to run
before website-src, since that is what exports the homepage including the
list they generate. The other way around, every freshly added post shows up on
the homepage one build late.
("website" :components
("website-posts-index" "website-posts-index-preview" "website-src"
"website-rss" "website-files" "website-static"))
Makefile#
Let's use make to avoid building and managing the website by hand.
OUT_DIR='www'
PUBLISH_FILE='publish-website.el'
PUBLISH_FUNC='(org-publish "website")'
WS_CMD=python -m http.server 12345 --bind localhost --directory $(OUT_DIR)
all:
rm -rf .cache www
emacs -Q --batch --load $(PUBLISH_FILE) --eval $(PUBLISH_FUNC)
rm -rf .cache
publish:
git commit
git push -u origin main
update:
make all
git add .
git status
run:
$(WS_CMD)
.PHONY: all mathjax publish update run
Whenever I want to work on a new item, like adding a new file or editing an
existing one, I switch to a new local branch in the git repository. I then
simply run make periodically while editing to visualize my changes. When I am
done, I merge the branch into my local main branch and run make update to
verify what will be committed and then make publish to push the changes on the
remote repository. Cherry on top, make run spawns a python web-server to view
the website locally at http://localhost:12345.
Deployment#
Once a commit is pushed, the GitHub Actions workflow takes over: it checks the repository out and hands the www/ directory to GitHub Pages as a deployment artifact.
on:
push:
branches: ["main"]
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
jobs:
deploy:
environment:
name: github-pages
url: ${{ steps.deployment.outputs.page_url }}
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Setup Pages
uses: actions/configure-pages@v5
- name: Upload artifact
uses: actions/upload-pages-artifact@v3
with:
path: 'www'
- name: Deploy to GitHub Pages
id: deployment
uses: actions/deploy-pages@v4
Note what this workflow does not do: there is no Emacs, no make, no build
step of any kind. It deploys a plain checkout, which is why www/ has to be
committed along with the sources – gitignoring the build output, the first
instinct for a generated directory, deploys an empty site. That choice has one
consequence worth knowing about: Org stamps every exported page with a build
date comment, which alone made all of them show up as modified after each
rebuild, so org-export-time-stamp-file is set to nil in publish-website.el.
The permissions: block is not optional either: without pages: write and
id-token: write, actions/deploy-pages fails outright.
The last piece worth mentioning is the CNAME file sitting at the root of the repository. It contains a single line – the domain name – which is how Pages knows to serve the site as www.skywhi.net instead of the default github.io address. The matching DNS records live at my registrar.
A few more functionalities#
MathJax#
MathJax' Org mode support lets me export to HTML math formulas written in LaTeX.
Here are a couple of examples using the inline $...$ and \( ... \)
delimiters as well as "proper" LaTeX environments.
If $a^2=b$ and \( b=2 \), then the solution must be either $$ a=+\sqrt{2} $$ or
\[ a=-\sqrt{2} \]
Look at this lonely equation I found:
\begin{equation}
y = x^2+1
\end{equation}
Which will be rendered like this:
If \(a^2=b\) and \( b=2 \), then the solution must be either \[ a=+\sqrt{2} \] or \[ a=-\sqrt{2} \]
Look at this lonely equation I found:
\begin{equation} y = x^2+1 \end{equation}
I install and update regularly MathJax via the Makefile and then enable it
in publish-website.el via the org-html-mathjax-options variable. Since I
want to be able to view the website offline – and since I would rather not
have my visitors fetch anything from a third party – I opted for a local
install.
Upstream ships around 20 MB: several output formats, Node entry points and a
speech locale for every supported language. A page here loads a small fraction
of that, so the target below pulls the release straight from the npm registry
(which, unlike a shallow clone of the default branch, is pinned to a version)
and keeps only the pieces that are actually served. The font package has a
version of its own, MATHJAX_FONT_VERSION, which merely follows
MATHJAX_VERSION for now: the two are released separately upstream and nothing
guarantees they stay in step.
MATHJAX_DIR=static/js/mathjax
MATHJAX_VERSION=4.1.3
MATHJAX_FONT=mathjax-modern
MATHJAX_FONT_VERSION=$(MATHJAX_VERSION)
NPM=https://registry.npmjs.org
TMP=$(MATHJAX_DIR).tmp
mathjax:
rm -rf $(MATHJAX_DIR) $(TMP)
mkdir -p $(TMP)/mathjax $(TMP)/font
curl -fsSL $(NPM)/mathjax/-/mathjax-$(MATHJAX_VERSION).tgz \
| tar xz -C $(TMP)/mathjax --strip-components=1
curl -fsSL $(NPM)/@mathjax/$(MATHJAX_FONT)-font/-/$(MATHJAX_FONT)-font-$(MATHJAX_FONT_VERSION).tgz \
| tar xz -C $(TMP)/font --strip-components=1
# [...] the rest of the target keeps only the files that are actually served.
Since version 4, MathJax no longer bundles its fonts: it resolves the font
package against loader.paths.fonts, which defaults to a jsdelivr CDN.
Vendoring the JavaScript is therefore not enough – the browser will still reach
out for roughly a megabyte of woff2 unless that path is pointed at the local
copy too. Org's stock MathJax template has no loader section, so I inject one:
;; MathJax resolves its font package against 'loader.paths.fonts', which
;; defaults to cdn.jsdelivr.net. The fonts are vendored (see the 'mathjax'
;; target in the Makefile), so that path has to be pointed at the local copy:
;; without it, every page carrying math quietly pulls ~1 MB of woff2 from a
;; third party -- on a site that otherwise makes no external request at all.
;; Org's stock template has no 'loader' section, hence the injection.
(defun skw-blog/mathjax-template-with-local-fonts (template)
"Return 'template' with a loader path pointing MathJax at the vendored fonts."
(let ((anchor "window.MathJax = {"))
(unless (string-match-p (regexp-quote anchor) template)
(error "MathJax template changed upstream: cannot inject the local font path"))
(replace-regexp-in-string
(regexp-quote anchor)
(concat anchor "\n loader: {paths: {fonts: '"
skw-blog/mathjax-dir "/fonts'}},")
template t t)))
(setq org-html-mathjax-options
`((path ,(concat skw-blog/mathjax-dir "/tex-mml-chtml.js"))
(scale 1.0) (align "center") (font "mathjax-modern") (overflow "overflow")
(tags "ams") (indent "0em") (multlinewidth "85%") (tagindent ".8em")
(tagside "right"))
org-html-mathjax-template
(skw-blog/mathjax-template-with-local-fonts org-html-mathjax-template))
RSS feed#
Generating an RSS feed to track when new posts are published is actually not that
complicated with ox-rss. Here are a couple of helper functions to generate a
"sitemap" file named rss.org which will list all the entries under
src/posts/ excluding the previously generated indexes. This file is then
converted to XML. The description field of each item is the content of the
preview tags used earlier.
;; RSS feed generation
(defun skw-blog/publish-to-rss (plist filename dir)
"Publish 'plist' when 'filename' corresponds to RSS feed Org-file to 'dir'."
(if (equal skw-blog/rss-filename (file-name-nondirectory filename))
;; Not `org-rss-publish-to-rss': it stamps a random :ID: on every headline
;; of rss.org and writes it back, so a plain "make all" dirties the source
;; tree for nothing -- <guid> is the permalink here anyway.
(org-publish-org-to 'rss filename
(concat "." (or (plist-get plist :rss-extension)
org-rss-extension))
plist dir)))
(defun skw-blog/format-rss-feed (title list)
"Generate a sitemap of posts that will be exported as an RSS feed. 'title' is
title of the RSS feed and 'list' the files to be included."
(concat "#+TITLE: " title "\n\n" (org-list-to-subtree list)))
(defun skw-blog/format-rss-feed-entry (entry style project)
"Format 'entry' for the posts RSS feed in given 'project'."
(let* ((title (org-publish-find-title entry project))
(link (concat (file-name-sans-extension entry) ".html"))
(pubdate (format-time-string (car org-time-stamp-formats)
(org-publish-find-date entry project)))
(preview (skw-blog/get-preview (concat (skw-blog/get-root-directory) "src/posts/" entry))))
(format "%s
:properties:
:rss_permalink: %s
:pubdate: %s
:end:
%s" title link pubdate preview)))
One wrinkle is worth mentioning, since it cost me a while to notice. The
obvious thing to call here is org-rss-publish-to-rss, but that function runs
org-icalendar-create-uid on the feed file before exporting it, which writes a
fresh random :ID: into every headline of rss.org and saves the file back
into src/. As rss.org is itself regenerated on each build, those IDs are new
every time and the file shows up as modified after every make – for IDs that
never even reach the feed, since org-rss-use-entry-url-as-guid is t and the
<guid> is simply the post permalink. Calling the exporter directly skips the
whole dance and leaves the source tree alone.
These functions are then called inside a new publishing rule.
("website-rss"
:author ,skw-blog/author
:auto-sitemap t
:base-directory ,(concat skw-blog/srcdir "posts")
:base-extension "org"
:description ,skw-blog/rss-description
:email ,skw-blog/email
:exclude ,(regexp-opt '("rss.org" "index.org" "index-no-preview.org"))
:html-link-home ,(concat skw-blog/upstream-url "/posts/")
:html-link-org-files-as-html t
:html-link-use-abs-url t
:publishing-directory ,skw-blog/outdir
:publishing-function skw-blog/publish-to-rss
:recursive nil
:rss-extension "xml"
:rss-feed-url ,(concat skw-blog/upstream-url "/rss.xml")
:rss-image-url ,(concat skw-blog/upstream-url "/static/img/profile.jpg")
:sitemap-filename ,skw-blog/rss-filename
:sitemap-format-entry skw-blog/format-rss-feed-entry
:sitemap-function skw-blog/format-rss-feed
:sitemap-sort-files anti-chronologically
:sitemap-title ,skw-blog/rss-feedname
:with-author t)
One entry deserves a word of explanation: :rss-feed-url. The feed is written
to the root of the site while its sources live under posts, so ox-rss
cannot work out the self-referencing <atom:link> from :html-link-home on its
own – it has to be spelled out, or readers end up subscribing to an address
that does not exist.
A feed nobody can find is not much use, so every exported page advertises it in
its <head> with a <link rel"alternate">= tag – the autodiscovery bit that
lets a browser or a feed reader pick the feed up from any page of the site.
That tag is skw-blog/rss-link, which website-src concatenates into its
:html-head next to the stylesheets and the favicon.
;; Feed autodiscovery, so readers can find the feed from any page
(setq skw-blog/rss-link
(concat "<link rel=\"alternate\" type=\"application/rss+xml\" title=\""
skw-blog/rss-feedname "\" href=\"/rss.xml\">\n"))
Heading permalinks#
Headings are made linkable via an anchor tag # with an org export filter.
First we need a way to detect if the export is for pure HTML content:
(defun skw-blog/html-page-backend-p (backend)
"Return non-nil when 'backend' exports whole HTML pages rather than feed items."
(and (org-export-derived-backend-p backend 'html)
(not (org-export-derived-backend-p backend 'rss))))
Only headings carrying a CUSTOM_ID get an anchor. Org derives every other id
from a counter that shifts as soon as the document changes, so such a link would
rot silently, hence the org[0-9a-f]+ test below to detect generated IDs.
(defun skw-blog/add-heading-anchor (headline backend info)
"Append a permalink anchor to the heading that opens 'headline'.
Only headings carrying a CUSTOM_ID get one: Org derives every other id from
a counter that shifts as soon as the document changes, so such a link would
rot silently."
(if (not (skw-blog/html-page-backend-p backend))
headline
(save-match-data
;; Children are filtered before their parent, so the first heading tag in
;; 'headline' is always the one this call is responsible for.
(if (not (string-match "<h\\([1-6]\\) id=\"\\([^\"]+\\)\">" headline))
headline
(let ((level (match-string 1 headline))
(id (match-string 2 headline))
(from (match-end 0)))
(if (string-match-p "\\`org[0-9a-f]+\\'" id)
headline
(let ((close (string-match (concat "</h" level ">") headline from)))
(if (null close)
headline
(concat (substring headline 0 close)
(format (concat "<a class=\"heading-anchor\" href=\"#%s\""
" aria-label=\"Permalink to this section\">#</a>")
id)
(substring headline close))))))))))
(add-to-list 'org-export-filter-headline-functions
#'skw-blog/add-heading-anchor)
Link previews and canonical URLs#
Org gives each page a <title> but nothing else: no canonical URL, and none of
the Open Graph properties that messaging apps, forums and social platforms read
to build a link preview. A page shared anywhere therefore rendered as a bare
URL, with no blurb and no image.
A filter on the final output injects the missing tags right before </head>.
The path a page is served at is derived from its source file, so nothing has to
be declared by hand – index.org files become directory URLs, everything else
keeps its name.
(defun skw-blog/page-path (input-file)
"Return the site-root-relative path at which 'input-file' is served."
(let* ((relative (file-relative-name (expand-file-name input-file)
skw-blog/srcdir))
(html (concat (file-name-sans-extension relative) ".html")))
(cond ((equal html "index.html") "/")
((string-suffix-p "/index.html" html)
(concat "/" (file-name-directory html)))
(t (concat "/" html)))))
(defun skw-blog/attribute-value (string)
"Return 'string' escaped for use as an HTML attribute value.
'org-html-encode-plain-text' leaves double quotes alone, which is fine for
text nodes and silently breaks an attribute the first time a title or a
description contains one."
(replace-regexp-in-string "\"" """ (org-html-encode-plain-text string) t t))
(defun skw-blog/page-metadata (info)
"Return the canonical, Open Graph and Twitter tags for the page in 'info'."
(let* ((input (plist-get info :input-file))
(path (and input (skw-blog/page-path input)))
(url (and path (concat skw-blog/upstream-url path)))
(title (skw-blog/attribute-value
(org-element-interpret-data (plist-get info :title))))
(description (skw-blog/attribute-value skw-blog/rss-description))
;; Posts and notes are articles; the homepage and the standing pages
;; around them are not.
(type (if (and path (string-match-p "\\`/\\(posts\\|notes\\)/." path))
"article"
"website"))
(tags '()))
(when url
(push (format "<link rel=\"canonical\" href=\"%s\">" url) tags)
(push (format "<meta property=\"og:url\" content=\"%s\">" url) tags))
(push (format "<meta property=\"og:type\" content=\"%s\">" type) tags)
(push (format "<meta property=\"og:site_name\" content=\"%s\">"
(skw-blog/attribute-value skw-blog/site-name))
tags)
(push (format "<meta property=\"og:title\" content=\"%s\">" title) tags)
(push (format "<meta property=\"og:description\" content=\"%s\">" description) tags)
(push (format "<meta property=\"og:image\" content=\"%s\">" skw-blog/og-image) tags)
(push "<meta name=\"twitter:card\" content=\"summary\">" tags)
;; The 404 is served for every unknown path. Indexing it would scatter
;; "No such file or directory" across search results.
(when (and input (equal (file-name-nondirectory input) "404.org"))
(push "<meta name=\"robots\" content=\"noindex\">" tags))
(concat (mapconcat #'identity (nreverse tags) "\n") "\n")))
(defun skw-blog/insert-page-metadata (output backend info)
"Insert the tags built by 'skw-blog/page-metadata' into 'output's head."
(if (not (skw-blog/html-page-backend-p backend))
output
(save-match-data
(if (not (string-match "</head>" output))
output
(replace-match (concat (skw-blog/page-metadata info) "</head>")
t t output)))))
(add-to-list 'org-export-filter-final-output-functions
#'skw-blog/insert-page-metadata)
The 404 gets a noindex, since it is served for every unknown path and having
"No such file or directory" scattered across search results helps nobody. And
the preview image is a variable of its own:
;; Shown when a page is shared on a platform that renders a link preview.
;; A 256x256 crop of the same picture as the favicon: the square shape suits
;; the summary card, and profile.jpg is 138x138, under the 144px floor below
;; which several platforms drop the image entirely.
(setq skw-blog/og-image
(concat skw-blog/upstream-url "/static/img/og-image.jpg"))
Resources#
This publishing setup is quite rudimentary. I wrote this note mostly for self reference in case I have to implement a similar project in the future. However, I still hope it can help hesitating people to give a try to org-publish and see if it fits their needs. Maybe existing users could also get a tip or two out of this. The full code is available on my github.
I would encourage anyone fiddling with this to have the official documentation right at hand: Org, Org-publish and Worg. Here are websites that inspired me for this adventure. I highly recommend that you check them out:
- Building a Emacs Org-Mode Blog by Thomas Ingram showcases a simple yet effective setup.
- Blogging using org-mode (and nothing else) by Dennis Ogbe offers a really neat explanation and setup, with Org-file pre-processing before publishing.
- A literate programming example by ryuslash.
Edit : Updated directory structure and rephrased some sections.
Edit : Update a lot of elements (sitemaps / indexes, RSS, build instructions, Github workflow, etc.).
Edit : Document the SEO files, the page metadata and the heading anchors, and resync the excerpts with the source files.