<?xml version="1.0" encoding="UTF-8" standalone="yes"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-us"><id>https://blog.pgxn.org/tags/api/</id><title>API</title><updated>2011-04-29T20:44:02Z</updated><link rel="self" type="application/atom+xml" href="https://blog.pgxn.org/tags/api/feed.xml"/><link rel="alternate" type="text/html" href="https://blog.pgxn.org/tags/api/"/><author><name>The PGXN Maintainers</name></author><generator uri="https://gohugo.io/" version="0.167.0">Hugo</generator><entry><id>https://blog.pgxn.org/post/5049235040</id><title type="html">About the Infrastructure: PGXN API</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/about-pgxn-api/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-04-29T20:44:02Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="infrastructure" label="Infrastructure"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="api-server" label="API Server"/><category scheme="https://blog.pgxn.org/tags" term="metadata" label="Metadata"/><category scheme="https://blog.pgxn.org/tags" term="rest" label="REST"/><category scheme="https://blog.pgxn.org/tags" term="http" label="Http"/><summary type="html"><![CDATA[<p>The second piece of the PGXN infrastructure, after <a href="https://blog.pgxn.org/post/4854707157/pgxn-manager" title="About the Infrastructure: PGXN Manager">PGXN Maanager</a>, is the
<a href="https://api.pgxn.org/">PGXN API Server</a>. I&rsquo;ve just finished the <a href="https://github.com/pgxn/pgxn-api/wiki">API documentation</a>, which covers
both the lightweight static file API provided by mirrors and the superset
provided by the API server. So now seems like a good time to talk about the
design of the API server and how it works.</p>
<p>At its core, the PGXN API server is just another mirror. It has an hourly cron
job that <code>rsync</code>s to the master mirror, updating the mirror. But then it
iterates over the <code>rsync</code> log and transforms some things. Here&rsquo;s what it does:</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>The second piece of the PGXN infrastructure, after <a href="https://blog.pgxn.org/post/4854707157/pgxn-manager" title="About the Infrastructure: PGXN Manager">PGXN Maanager</a>, is the
<a href="https://api.pgxn.org/">PGXN API Server</a>. I&rsquo;ve just finished the <a href="https://github.com/pgxn/pgxn-api/wiki">API documentation</a>, which covers
both the lightweight static file API provided by mirrors and the superset
provided by the API server. So now seems like a good time to talk about the
design of the API server and how it works.</p>
<p>At its core, the PGXN API server is just another mirror. It has an hourly cron
job that <code>rsync</code>s to the master mirror, updating the mirror. But then it
iterates over the <code>rsync</code> log and transforms some things. Here&rsquo;s what it does:</p>
<ul>
<li>Unpacks each distribution in a directory for browsing. Here, for example, is
where one can <a href="https://api.pgxn.org/src/semver/semver-0.2.1/">browse the semver 0.2.1</a> sources.</li>
<li>Searches for a <code>README</code> file and any files recognized by <a href="https://search.cpan.org/perldoc?Text::Markup">Text::Markup</a> and
converts them to sanitized HTML with a table of contents. Such files can
then be used to display the <code>README</code> on the <a href="https://pgxn.org/dist/semver/" title="semver distribution page">distribution page</a> and to
display <a href="https://pgxn.org/dist/semver/doc/semver.html" title="semver documentation">individual documentation files</a>.</li>
<li>Merges the distribution metadata file with the latest stable release
<code>META.json</code> generated by PGXN Manager. For example, as of this writing, the
API server&rsquo;s <a href="https://api.pgxn.org/dist/semver/0.2.1/META.json">semver 0.2.1 <code>META.json</code></a> and the unversioned <a href="https://api.pgxn.org/dist/semver.json">semver.json</a>
are identical. Effectively, this format has all the metadata from the
<code>META.json</code> as well as a list of all releases of the distribution from the
<code>semver.json</code>. This is useful for displaying all the data on the
<a href="https://pgxn.org/dist/semver/" title="semver distribution page">distribution page</a> by fetching the data in a single API request.</li>
<li>Updates all other versions of the <code>META.json</code> file. For example, if you look
at the <a href="https://api.pgxn.org/dist/semver/0.2.0/META.json">semver 0.0.0 <code>META.json</code></a>, you&rsquo;ll see that it includes 0.2.1 in its
list of releases, even though 0.2.1 was released after 0.2.0. This allows
<a href="https://pgxn.org/dist/semver/0.2.0/">semver 0.2.0</a> page on the main site to have a select list of version to
choose from, including versions released later, with a single API request.</li>
<li>Adds additional metadata to the extension JSON file for all extensions in
the distribution. The added data includes release dates for the list all
distributions providing the extension, as well as an abstract and doc path
for the latest stable release. To see the differences, compare the <a href="https://api.pgxn.org/mirror/extension/semver.json">mirror
<code>semver.json</code></a> to the <a href="https://api.pgxn.org/extension/semver.json">API <code>semver.json</code></a>.</li>
<li>Adds an abstract for each distribution listed in the user&rsquo;s JSON file and
all tag JSON files. Compare, for example, the <a href="https://api.pgxn.org/mirror/user/theory.json">mirror <code>theory.json</code></a> to the
<a href="https://api.pgxn.org/user/theory.json">API <code>theory.json</code></a> and the <a href="https://api.pgxn.org/mirror/tag/data%20types.json">mirror <code>data types.json</code></a> to the <a href="https://api.pgxn.org/tag/data%20types.json">API
<code>data types.json</code></a>. This allows the <a href="https://pgxn.org/user/theory">user page</a> and <a href="https://pgxn.org/tag/data%20types/">tag pages</a> to include
the abstract in the list of distributions released by the user or associated
with a tag.</li>
<li>Adds records to a <a href="https://incubator.apache.org/lucy/">Lucy</a>-powered full text search index.</li>
</ul>
<p>All of this merging stuff came out of my thinking following the discussion of
the <a href="https://blog.pgxn.org/post/3099288750/pgxn-api-rfc">PGXN API RFC</a>. The decision to use <a href="https://incubator.apache.org/lucy/">Lucy</a> instead of PostgreSQL&rsquo;s
<a href="https://www.postgresql.org/docs/current/static/textsearch.html">full-text search</a> followed rather naturally from this, as I quickly realized
that there was no other driving need for a relational database behind the API
at all. The <em>only</em> dynamic API is the <a href="https://github.com/pgxn/pgxn-api/wiki/search-api">search API</a>. Everything else is just
static files. And given the <a href="https://www.depesz.com/index.php/2010/10/17/why-im-not-fan-of-tsearch-2/">performance issues</a> of in-database search, as
well as the desire to have fewer outside dependencies, made the decision a
natural one.</p>
<p>Beyond the syncing, there is a very simple web server providing the HTTP REST
interface to the static JSON files and the full-text search. That&rsquo;s it,
really. The API server is really just another mirror on steroids. The nice
thing is that it allows an interface, such as <a href="https://search.cpan.org/perldoc?WWW::PGXN">WWW::PGXN</a> or the new <a href="https://blog.pgxn.org/post/5026314153/writing-a-client-for-pgxn">PGXN
client</a> to work with either interface, just failing gracefully when API server
APIs are unavailable.</p>
<p>If you want to learn more about the specifics of the REST API, the <a href="https://github.com/pgxn/pgxn-api/wiki">API
documentation</a> has <em>all</em> the details. Really, it&rsquo;s quite comprehensive!</p>
<p>I actually consider the API to be 1.0-complete at this point, unlike PGXN
Manager. The only thing I want to add is <a href="https://en.wikipedia.org/wiki/JSONP">JSONP</a> support for static JSON files
(right now it&rsquo;s only for search results) and might tweak a few things here and
there, but otherwise I think it&rsquo;s in pretty good shape.</p>
<p>Longer term, though, it might be worthwhile to add some other features to
enhance the value of PGXN overall. Some ideas:</p>
<ul>
<li>Distribution and/or extension ratings (reviews, Like/Dislike, stars, or
something).</li>
<li>Diffs to compare changes between versions.</li>
<li>A test reporting infrastructure with result matrices (á la <a href="https://www.cpantesters.org/">CPAN Testers</a>.</li>
</ul>
<p>But I think we need to build up some momentum on the foundation that&rsquo;s in
place. Have you submitted your extensions, yet?</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/4970682941</id><title type="html">PGXN API Docs Published</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/api-docs-published/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-04-27T00:27:55Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="documentation" label="Documentation"/><category scheme="https://blog.pgxn.org/tags" term="docs" label="Docs"/><category scheme="https://blog.pgxn.org/tags" term="api-docs" label="API Docs"/><category scheme="https://blog.pgxn.org/tags" term="zip" label="Zip"/><category scheme="https://blog.pgxn.org/tags" term="client" label="Client"/><category scheme="https://blog.pgxn.org/tags" term="daniele-varrazzo" label="Daniele Varrazzo"/><summary type="html"><![CDATA[<p>The <a href="https://github.com/pgxn/pgxn-api/wiki">PGXN API Documentation</a> is up! I&rsquo;ve just finished writing the docs for
the lightweight REST API provided by all PGXN mirrors. It also documents how
the same APIs differ when provided by the <a href="https://api.pgxn.org/">API server</a>. Their are four more
documents to write still, APIs provided by the API server but not the mirrors
(including full-text search); I should be able to get those done tomorrow.</p>
<p>So if you&rsquo;re interested in use the API for various things, please <a href="https://github.com/pgxn/pgxn-api/wiki">have a
look</a>! I&rsquo;d appreciate any feedback or corrections.</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>The <a href="https://github.com/pgxn/pgxn-api/wiki">PGXN API Documentation</a> is up! I&rsquo;ve just finished writing the docs for
the lightweight REST API provided by all PGXN mirrors. It also documents how
the same APIs differ when provided by the <a href="https://api.pgxn.org/">API server</a>. Their are four more
documents to write still, APIs provided by the API server but not the mirrors
(including full-text search); I should be able to get those done tomorrow.</p>
<p>So if you&rsquo;re interested in use the API for various things, please <a href="https://github.com/pgxn/pgxn-api/wiki">have a
look</a>! I&rsquo;d appreciate any feedback or corrections.</p>
<p>Speaking of users of the API, <a href="https://profiles.google.com/daniele.varrazzo/about">Daniele Varrazzo</a> has started writing a <a href="https://github.com/dvarrazzo/pgxn-client/">PGXN
client app</a> in Python. This is so awesome! It makes me happy that people are
able to just get going on this. And Daniele did it before I&rsquo;d written the
docs, just from reading the PGXN source code. Nice!</p>
<p>Anyway, getting these docs written was the last barrier I had to releasing
various pieces of PGXN on CPAN and generally being able to get on with my
life. I wanted to write these docs first so that I would have a little more
freedom to change things if something struct me as especially stupid.
Fortunately, as I&rsquo;ve written the <a href="https://pgxn.org/">main site</a> as a thin client for the API,
most of the lameness has already been excised. I think the only change I&rsquo;d
like to make is to change the <code>{char}</code> URI template variable for the
<a href="https://github.com/pgxn/pgxn-api/wiki/userlist-api"><code>userlist</code> API</a> to <code>{letter}</code>, because only lowercased ASCII letters a-z are
allowed. It&rsquo;s more accurate.</p>
<p>Another thing I think I&rsquo;ll change is the file name suffix used for the
distribution download files. Currently it&rsquo;s <code>.pgz</code>, but after <a href="https://blog.pgxn.org/post/4854707157/pgxn-manager">Daniele
complained</a> about it, I <a href="https://groups.google.com/group/pgxn-users/browse_thread/thread/1209538be03be8a4">asked around</a> and the concensus seems to be to change
it to <code>.zip</code>. I&rsquo;ll likely do that tomorrow, too. Fortunately it&rsquo;s pretty easy:
Just edit the configuration file and rename the files on the master mirror. At
least it <em>should</em> be that easy!</p>
<p>There was one other thing in the APIs the felt a bit silly, but I don&rsquo;t
remember what it was, so screw it.</p>
<p>Anyway, feedback on the API docs would be greatly appreciated. And &ndash; <em>get
hacking!</em>.</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/4601750614</id><title type="html">It&amp;rsquo;s Finally Up: The New PGXN Site!</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/new-pgxn-site/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-04-14T06:32:06Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="site" label="Site"/><category scheme="https://blog.pgxn.org/tags" term="launch" label="Launch"/><category scheme="https://blog.pgxn.org/tags" term="final" label="Final"/><category scheme="https://blog.pgxn.org/tags" term="awesome" label="Awesome"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><summary type="html"><![CDATA[<p>I&rsquo;m pleased to announce that the new <a href="https://pgxn.org/">PGXN SITE</a> went live last night. Some of
the things it does:</p>
<ul>
<li>Full text search of all documentation, distribution metadata and <code>README</code>s,
extensions metadata, and users</li>
<li>list of the 56 most recent uploads</li>
<li>Find users by first letter of nickname (mostly so search spiders can drill
down to all site content)</li>
<li>A tag cloud with the 56 most commonly-used tags (they&rsquo;re all the same size
and color right now because each one is used only once at the moment)</li>
<li>Pages for users (<a href="https://pgxn.org/user/alexk">example</a>)</li>
<li>Pages for distributions, with inlined <code>README</code> (<a href="https://pgxn.org/dist/explanation/">example</a>)</li>
<li>Browse unzipped distributions (<a href="https://api.pgxn.org/src/semver/semver-0.2.0/">example</a>)</li>
<li>Pages for documentation (<a href="https://pgxn.org/dist/explanation/doc/explanation.html">example</a>)</li>
<li>Permalinks for extensions <a href="https://pgxn.org/extension/pair">example</a>)</li>
</ul>
<p>The entire site is backed by the <a href="https://api.pgxn.org/">API Server</a>, which is mainly composed of
static JSON and HTML files read by the site back end from the local file
system. So the site is quite fast. Please browse around, let me know what you
think, and please <a href="https://github.com/pgxn/pgxn-site/issues">report any bugs</a> you find.</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I&rsquo;m pleased to announce that the new <a href="https://pgxn.org/">PGXN SITE</a> went live last night. Some of
the things it does:</p>
<ul>
<li>Full text search of all documentation, distribution metadata and <code>README</code>s,
extensions metadata, and users</li>
<li>list of the 56 most recent uploads</li>
<li>Find users by first letter of nickname (mostly so search spiders can drill
down to all site content)</li>
<li>A tag cloud with the 56 most commonly-used tags (they&rsquo;re all the same size
and color right now because each one is used only once at the moment)</li>
<li>Pages for users (<a href="https://pgxn.org/user/alexk">example</a>)</li>
<li>Pages for distributions, with inlined <code>README</code> (<a href="https://pgxn.org/dist/explanation/">example</a>)</li>
<li>Browse unzipped distributions (<a href="https://api.pgxn.org/src/semver/semver-0.2.0/">example</a>)</li>
<li>Pages for documentation (<a href="https://pgxn.org/dist/explanation/doc/explanation.html">example</a>)</li>
<li>Permalinks for extensions <a href="https://pgxn.org/extension/pair">example</a>)</li>
</ul>
<p>The entire site is backed by the <a href="https://api.pgxn.org/">API Server</a>, which is mainly composed of
static JSON and HTML files read by the site back end from the local file
system. So the site is quite fast. Please browse around, let me know what you
think, and please <a href="https://github.com/pgxn/pgxn-site/issues">report any bugs</a> you find.</p>
<p>There are a few more things left to do on my list:</p>
<ul>
<li>Add pages for users who don&rsquo;t yet have distributions on the network.</li>
<li>Convert the localization interface to <code>gettext</code> and solicit translations.</li>
<li>Document the API; might lead to some changes as I run across
inconsistencies.</li>
<li>Add an Atom feed of recent releases.</li>
<li>Add aliases for <code>$nickname@pgxn.org</code> addresses, to forward to personal
addresses; this way we can just publish pgxn.org addresses.</li>
<li>Write documentation on how to optimize distributions for optimal exposure on
PGXN.</li>
<li>Consider changing the Permalink URL to something shorter. Suggestions?</li>
</ul>
<p>If you&rsquo;d liked to help out with any of these tasks, just let me know. Better
yet, grab the source from <a href="https://github.org/pgxn/">GitHub</a> and just hack!</p>
<p>Oh, and speaking of the source, I would <em>really</em> appreciate a code review or
four. Although I have solicited advice on key questions fro people more
knowledgeable than I (thanks Aristotle, Miyagawa, Andreas, Graham), the code
is entirely my own. More eyes will be a huge help!</p>
<p>Tomorrow I&rsquo;ll blog a bit about the architecture for the network. I&rsquo;m quite
happy with it.</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/3641963548</id><title type="html">Question: Ajax and Search Engine Indexing</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/ajax-search-indexing/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-03-04T20:20:08Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="ajax" label="Ajax"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="search-engine" label="Search Engine"/><category scheme="https://blog.pgxn.org/tags" term="indexing" label="Indexing"/><category scheme="https://blog.pgxn.org/tags" term="indexability" label="Indexability"/><category scheme="https://blog.pgxn.org/tags" term="response-code" label="Response Code"/><category scheme="https://blog.pgxn.org/tags" term="not-found" label="Not Found"/><summary type="html"><![CDATA[I&rsquo;ve started working on the main (search) site in earnest now. The basic
layout is done, and I&rsquo;m working on the distribution view (<a href="https://theory.github.com/pgxn/pgtapdist.html">mockup</a>). My
thinking so far has been that I would simply serve a page that requested, say,
<code>/dist/pgTAP/</code>, and that page would use Ajax requests to fetch the data from
the <a href="https://api.pgxn.org/">API server</a> and display stuff. I think this will work pretty well except
for one thing: 404s.]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I&rsquo;ve started working on the main (search) site in earnest now. The basic
layout is done, and I&rsquo;m working on the distribution view (<a href="https://theory.github.com/pgxn/pgtapdist.html">mockup</a>). My
thinking so far has been that I would simply serve a page that requested, say,
<code>/dist/pgTAP/</code>, and that page would use Ajax requests to fetch the data from
the <a href="https://api.pgxn.org/">API server</a> and display stuff. I think this will work pretty well except
for one thing: 404s.</p>
<p>That is, if you request <code>/dist/nonexistent/</code>, then it will load a page with
the HTTP status code <code>200 OK</code>, but then, when the Ajax request 404s, it will
show a &ldquo;Not found&rdquo; error message. That&rsquo;s all well and good, but I&rsquo;m wondering
about the impact of two things:</p>
<ol>
<li>
<p>Since the page itself won&rsquo;t 404, search engines might index links to
nonexistent extensions. Of course, bad links won&rsquo;t be <em>that</em> common, but
of course they do happen and then tend to live forever.</p>
</li>
<li>
<p>If the search site uses Ajax to fetch the contents of a page via JSON (or,
for documentation, as an HTML document it will put into a div), will the
full content be properly indexed by search engines?</p>
</li>
</ol>
<p>So these are serious questions, in my mind. Do we loose good search engine
indelibility when we load content dynamically?</p>
<p>Of course, I can instead write it so that the back end fetches stuff from the
API server (and perhaps directly from the file system) and get &lsquo;round these
issues, but then it&rsquo;s less of a cool example of the use of the API server.</p>
<p>What do you think? Good advice much appreciated!</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/3099288750</id><title type="html">PGXN API RFC</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/pgxn-api-rfc/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-02-04T04:05:33Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="search" label="Search"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="mirror" label="Mirror"/><category scheme="https://blog.pgxn.org/tags" term="json" label="JSON"/><category scheme="https://blog.pgxn.org/tags" term="cpan" label="CPAN"/><category scheme="https://blog.pgxn.org/tags" term="metacpan" label="metacpan"/><category scheme="https://blog.pgxn.org/tags" term="javascript" label="JavaScript"/><category scheme="https://blog.pgxn.org/tags" term="application" label="Application"/><summary type="html"><![CDATA[<p>Things slowed up a bit over the last couple of months, I admit. There are any
number of reasons for that, not the least were the intrusion of the holidays
and a <a href="https://www.designsceneapp.com/">little side project</a> I&rsquo;ve been hacking on after-hours (and sometimes
during-hours). But I&rsquo;m ramping things up again now and need <em>your</em> feedback on
my current plans. Here&rsquo;s what I&rsquo;m working on: the search site.</p>
<h2 id="search-sites-and-apis">Search Sites and APIs</h2>
<p>Well, sort of. First of all, I&rsquo;ve decided that the &ldquo;search site&rdquo; should not be
a separate thing. The <a href="https://pgxn.org/">main site</a> will be the search site. This is following
the example of <a href="https://openjsan.org">JSAN</a>, as well as feedback from <a href="https://search.cpan.org/~gbarr/">Graham Barr</a>, who created and
maintains <a href="https://search.cpan.org/">CPAN Search</a>. Apparently people are often confused that
<a href="https://search.cpan.org/">search.cpan.org</a> is separate from <a href="https://www.cpan.org">www.cpan.org</a>. No point in
adding in confusion from the beginning. And besides, now that the PGXN
fund-raising <a href="https://blog.pgxn.org/post/2063677299/goooooooaaaal">is over</a>, I don&rsquo;t know what else would go on the home page.</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>Things slowed up a bit over the last couple of months, I admit. There are any
number of reasons for that, not the least were the intrusion of the holidays
and a <a href="https://www.designsceneapp.com/">little side project</a> I&rsquo;ve been hacking on after-hours (and sometimes
during-hours). But I&rsquo;m ramping things up again now and need <em>your</em> feedback on
my current plans. Here&rsquo;s what I&rsquo;m working on: the search site.</p>
<h2 id="search-sites-and-apis">Search Sites and APIs</h2>
<p>Well, sort of. First of all, I&rsquo;ve decided that the &ldquo;search site&rdquo; should not be
a separate thing. The <a href="https://pgxn.org/">main site</a> will be the search site. This is following
the example of <a href="https://openjsan.org">JSAN</a>, as well as feedback from <a href="https://search.cpan.org/~gbarr/">Graham Barr</a>, who created and
maintains <a href="https://search.cpan.org/">CPAN Search</a>. Apparently people are often confused that
<a href="https://search.cpan.org/">search.cpan.org</a> is separate from <a href="https://www.cpan.org">www.cpan.org</a>. No point in
adding in confusion from the beginning. And besides, now that the PGXN
fund-raising <a href="https://blog.pgxn.org/post/2063677299/goooooooaaaal">is over</a>, I don&rsquo;t know what else would go on the home page.</p>
<p>The other thing that&rsquo;s happened is, just as I was getting my butt in gear on
this stuff, a new CPAN search site came to my attention, <a href="https://search.metacpan.org/">μετα CPAN</a>. This is
an interesting project. What they did instead of creating a monolithic HTML
search site is to create a <a href="https://github.com/CPAN-API/cpan-api/wiki/API-docs">simple API</a> that serves nothing but JSON. It has
search and displays metadata for CPAN objects (distributions, maintainers,
modules, etc.). The search site, then, is not really a site at all, but a pure
JavaScript application. Once you load it, it just uses the API server to get
all the data. There are a few tricks server-side to proxy the API server so as
to avoid cross-site scripting issues. But otherwise it just works in the
browser.</p>
<p>Now I&rsquo;m not sure I&rsquo;ll do the same thing, exactly, but there&rsquo;s a lot of appeal
in creating a RESTful API server that&rsquo;s independent of the search site, and
then building the search site to use it. It also has the advantage of being
useful for other projects to just use. Want to create a PGXN search widget for
your blog? Yeah, there&rsquo;s an API for that.</p>
<h2 id="a-super-restful-directory">A Super RESTful Directory</h2>
<p>Of course, thanks to the &ldquo;RESTful Directory&rdquo; design for the mirrors (described
<a href="https://blog.pgxn.org/post/988613682/restful-directory" title="A RESTful Directory">here</a> and revised <a href="https://blog.pgxn.org/post/1138292188/arch-and-extension-json" title="Architecture, New Extension JSON Format">here</a>), any mirror is a lightweight API already.
There&rsquo;s a <em>lot</em> of metadata one can get just from the static JSON files it
generates. The design is flexible&ndash;but designed with a command-line client in
mind. As such, many commands executed in a command-line client would likely
requires multiple requests to a mirror. For example:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-console" data-lang="console"><span class="line"><span class="cl"><span class="gp">&gt;</span> install pgtap
</span></span></code></pre></div><p>This would request <code>/by/extension/pgtap.json</code> from the server. It would then
parse that file and see that the latest stable version of pgTAP is in the
distribution &ldquo;pgTAP&rdquo; at version &ldquo;0.25.0&rdquo;. So it would then download
<code>/dist/pgTAP-0.25.0.pgz</code> to install.</p>
<p>This is great for a command-line client, but wouldn&rsquo;t be so great for a search
site to be responsive. Ideally, a site should send a single request to get all
the data it needs for a particular page.</p>
<p>So here&rsquo;s what I&rsquo;m thinking for a PGXN API server: It will offer a superset of
the functionality of any other PGXN mirror. That is, all the JSON files in a
mirror will be present, but many of them will have more information than they
would on other mirrors. And then, of course, there will be other URIs to offer
additional API calls.</p>
<h3 id="details">Details</h3>
<p>So what does that look like? Let&rsquo;s take the pgTAP distribution, which I
released on PGXN earlier this week. To find the pgTAP distribution, one
requests:</p>
  <blockquote>
    <p><a href="https://master.pgxn.org/by/dist/pgTAP.json"><code>/by/dist/pgTAP.json</code></a></p>
  </blockquote>
<p>From that, one can see that the latest table release is 0.25.0, and so one can
then request</p>
  <blockquote>
    <p><a href="https://master.pgxn.org/dist/pgTAP/pgTAP-0.25.0.json"><code>/dist/pgTAP/pgTAP-0.25.0.json</code></a></p>
  </blockquote>
<p>to get all the metadata for that particular release. What I propose, to avoid
the two requests, is to include the contents of the second file in the first.
That would then have all the data necessary to generate the <a href="https://theory.github.com/pgxn/pgtapdist.html">pgTAP
distribution page</a> on the PGXN site.</p>
<p>The API would offer similar supersets of data for the <a href="https://master.pgxn.org/by/extension/pgtap.json">extension</a> , <a href="https://master.pgxn.org/by/owner/theory.json">owner</a> ,
and <a href="https://master.pgxn.org/by/tag/testing.json">tag</a> metadata files, to have the data necessary for the design of the
corresponding <a href="https://theory.github.com/pgxn/pgtap.html">extension</a>, <a href="https://theory.github.com/pgxn/theory.html">owner</a> and tag layouts of the site.</p>
<h3 id="additional-resources">Additional Resources</h3>
<p>In addition to adding metadata to the existing mirrored JSON files, there
would be other resources available for request from the API server. They would
include:</p>
<ul>
<li>
<p>Extension Documentation. Each distribution may include documentation for
included extensions in the <code>doc</code> subdirectory. These will go under the
directory for a specific distribution such as
<code>/dist/pgTAP/pgTAP-0.35.0/doc/pgtap.html</code>. The latest version of each
document would also be available under <code>/by/extension</code>, as in
<code>/by/extension/pgtap.html</code>. This requires that the documentation file have
the same base name as the extension file itself.</p>
</li>
<li>
<p>Other documentation. I&rsquo;d like to support arbitrary documentation, such as
for included binary executables, HOWTOs, etc. The canonical copies will go
under the versioned distribution URL, of course, but I&rsquo;m not sure about
permalinks. That might require an extension of the <a href="https://pgxn.org/meta/spec.txt">Meta Spec</a>; I haven&rsquo;t
quite figured that out, yet.</p>
</li>
<li>
<p>Source code. There will be an interface to browse an unpacked copy of any
distribution as plain text. This will be under <code>/src</code>, as in
<code>/src/pgTAP/pgTAP-0.35.0/</code>.</p>
</li>
</ul>
<h3 id="search-api">Search API</h3>
<p>Of course. This is the big one, really. I think it makes sense to have the
<code>/by</code> URI respond to search requests. Thus, a request for</p>
<pre tabindex="0"><code>/by?q=testing
</code></pre><p>would search everything. If you only want to search a certain category of
object, you&rsquo;d hit the appropriate URI:</p>
<pre tabindex="0"><code>/by/dist?q=tap
/by/owner?q=clochard
/by/tag?q=test
/by/extension?q=gis
</code></pre><p>The nice thing about this is that it retains the existing entity URLs. The
directory level determines which entities you get.</p>
<h2 id="your-thoughts">Your Thoughts?</h2>
<p>So that&rsquo;s my thinking on the search API. I&rsquo;m going to start hacking on it in
earnest tomorrow, and perhaps next week I can get a very early version out
(basically just another mirror to start with).</p>
<p>But what do you think? Seem like a sane approach? Am I missing anything
obvious or doing anything clearly stupid? Please let me know in the comments!</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/1082188310</id><title type="html">Status Update: DB API, Extension Versions RFC</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2010/db-status-update/"/><updated>2026-10-07T16:13:48Z</updated><published>2010-09-07T18:46:38Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="status-update" label="Status Update"/><category scheme="https://blog.pgxn.org/tags" term="database" label="Database"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="documentation" label="Documentation"/><category scheme="https://blog.pgxn.org/tags" term="extensions" label="Extensions"/><category scheme="https://blog.pgxn.org/tags" term="versions" label="Versions"/><category scheme="https://blog.pgxn.org/tags" term="rfc" label="RFC"/><summary type="html"><![CDATA[I spent most of the time I had to work on PGXN the last two weeks creating the
database for <a href="https://github.com/theory/pgxn-manager/">PGXN Manager</a>. I had estimated 24 hours of work to design the
database. So far I&rsquo;ve logged 34 hours. But I think that will come out in the
wash, really, because I did a lot of work to generate JSON from database
functions. This is code I had originally expected to do in the app layer. I&rsquo;m
starting work on that this week, with an estimated 40 hours. Hope I can do it
in 30. :-)]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I spent most of the time I had to work on PGXN the last two weeks creating the
database for <a href="https://github.com/theory/pgxn-manager/">PGXN Manager</a>. I had estimated 24 hours of work to design the
database. So far I&rsquo;ve logged 34 hours. But I think that will come out in the
wash, really, because I did a lot of work to generate JSON from database
functions. This is code I had originally expected to do in the app layer. I&rsquo;m
starting work on that this week, with an estimated 40 hours. Hope I can do it
in 30. :-)</p>
<p>I&rsquo;m pretty happy with how the database API is turning out. See the nifty
<a href="https://github.com/theory/pgxn-manager/wiki/DB-API">database API documentation</a> I whipped up using <a href="https://github.com/theory/pgxn-manager/blob/master/bin/gendoc">gendoc</a> and embedded
<a href="https://fletcherpenney.net/multimarkdown/users_guide/multimarkdown_syntax_guide/">MultiMarkdown</a>. There are still a few tweaks I need to make. I just checked
in a change to eliminate all the <code>ALIAS</code>es I <a href="https://blog.pgxn.org/post/1053165383/alias-in-vogue">blogged about</a> last week. Turns
out that one can just <a href="https://archives.postgresql.org/pgsql-hackers/2010-09/msg00404.php">function-name-qualify</a> the function parameter names.
Who knew? I sure didn&rsquo;t.</p>
<p>But I do have a question for you, dear readers. Right now, the database
requires that extensions have unique version numbers in every distribution. So
if, for example, you uploaded the distribution foo 1.2.2 with the extension
bar 1.2.2, when you next uploaded foo 1.2.3, bar could not be 1.2.2. I
designed this this way by following my own practice of always incrementing the
version numbers of all modules included in my <a href="https://search.cpan.org/~dwheeler/">CPAN distributions</a>. So all
modules have unique version numbers, with never a duplicate.</p>
<p>However, CPAN itself only cares that distributions have unique version
numbers. It doesn&rsquo;t care about the version numbers of included modules. See,
for example, <a href="https://search.cpan.org/dist/HTML-Mason/">HTML::Mason</a>. Note that the various included modules have all
sorts of different version numbers, and many have no version numbers at all.
So while the core module has a version number (1.45 at the time of this
writing), if you click to look at older versions, say <a href="https://search.cpan.org/~drolsky/HTML-Mason-1.44/">Mason 1.44</a>, you&rsquo;ll see
that no other modules have their version numbers changed.</p>
<p>I don&rsquo;t think I want to allow extensions without version numbers on PGXN.
That&rsquo;s been a recipe for many annoyances with CPAN. But maybe I should allow
extension version numbers to repeat? The upside is less maintenance for
multi-extension distribution authors. The downside is potentially less
accuracy in the determination of prerequisites.</p>
<p>For example, say that we have two versions of distribution foo:</p>
<pre tabindex="0"><code>  -----------------------------------------------------------------------------
  Distribution                           Extensions
  -------------------------------------- --------------------------------------
  foo 1.2.2                              foo 1.2.2\
                                         bar 1.2.2
<h2 id="bar-122">foo 1.2.3                              foo 1.2.3<br>
bar 1.2.2</h2>
<p></code></pre><p>Note how the version number of extension bar has not changed between releases.
But if I was a user of both foo and bar, and I specified that I required bar
1.2.2 but was actually using stuff in foo 1.2.3, I could end up with errors
because I would only have foo 1.2.3.</p></p>
<p>This is an issue periodically faced by Perl hackers. I might require
HTML::Mason::ApacheHandler 1.69 but really what I need is HTML::Mason 1.44. Of
course, as a Perl hacker, it&rsquo;s my responsibility to make sure I require
exactly what I need, but sometimes it&rsquo;s not clear what&rsquo;s the primary module in
a distribution. And if I don&rsquo;t realize that the version number of
HTML::Mason::ApacheHandler doesn&rsquo;t change with every release, I might make
mistakes.</p>
<p>Maybe that&rsquo;s okay. Maybe PGXN should be less rigid like this. But I&rsquo;m leaning
toward keeping things as they are and requiring new versions of every
extension in every upload of a distribution. It&rsquo;s a bit tougher for those
(few?) hackers who will create multi-extension distributions, but probably
more reliable for everyone else.</p>
<p>What do you think?</p>
<p>Anyway, I&rsquo;m getting to work on the Web app this week while this issue
percolates. Been studying up on <a href="https://plackperl.org/">Plack</a>. So nice!</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/988613682</id><title type="html">A RESTful Directory</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2010/restful-directory/"/><updated>2026-10-07T16:13:48Z</updated><published>2010-08-21T18:52:00Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="mirror" label="Mirror"/><category scheme="https://blog.pgxn.org/tags" term="directory" label="Directory"/><category scheme="https://blog.pgxn.org/tags" term="structure" label="Structure"/><category scheme="https://blog.pgxn.org/tags" term="meta" label="Meta"/><category scheme="https://blog.pgxn.org/tags" term="metadata" label="Metadata"/><category scheme="https://blog.pgxn.org/tags" term="json" label="JSON"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="scalability" label="Scalability"/><category scheme="https://blog.pgxn.org/tags" term="rest" label="REST"/><summary type="html"><![CDATA[Following my post outlining a possible network <a href="https://blog.pgxn.org/post/954535657/thoughts-on-the-network-directory-structure">directory structure</a>,
<a href="https://plasmasturm.org/">Aristotle Pagaltzis</a> saw fit to bug me via email about a different approach.
I couldn&rsquo;t understand WTF he was talking about until today. Then it lit my
brain on fire. As a result, I now think that there is a much better way to
organize the metadata files for the PGXN &ndash; one that happens not to include
any symbolic links (which is something that <a href="https://search.cpan.org/~andk/">Andreas König</a> has been flagging,
via email, as a possible bottleneck).]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>Following my post outlining a possible network <a href="https://blog.pgxn.org/post/954535657/thoughts-on-the-network-directory-structure">directory structure</a>,
<a href="https://plasmasturm.org/">Aristotle Pagaltzis</a> saw fit to bug me via email about a different approach.
I couldn&rsquo;t understand WTF he was talking about until today. Then it lit my
brain on fire. As a result, I now think that there is a much better way to
organize the metadata files for the PGXN &ndash; one that happens not to include
any symbolic links (which is something that <a href="https://search.cpan.org/~andk/">Andreas König</a> has been flagging,
via email, as a possible bottleneck).</p>
<p>First, the <code>/dist</code> directory will be the same as before. Releases of pgTAP
would be in:</p>
<pre tabindex="0"><code>dist/p/pg/pgtap/pgtap-0.23.pgz
dist/p/pg/pgtap/pgtap-0.23.json
dist/p/pg/pgtap/pgtap-0.23.readme
dist/p/pg/pgtap/pgtap-0.24.pgz
dist/p/pg/pgtap/pgtap-0.24.json
dist/p/pg/pgtap/pgtap-0.24.readme
dist/p/pg/pgtap/pgtap-0.25.pgz
dist/p/pg/pgtap/pgtap-0.25.json
dist/p/pg/pgtap/pgtap-0.25.readme
</code></pre><p>The only change is that the <code>pgtap.json</code> symlink is gone.</p>
<p>Now, the new stuff. In the root directory will be a file, <code>index.json</code>, that
contains templates for URIs. It will look something like this:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;dist&#34;</span><span class="p">:</span>   <span class="s2">&#34;/dist/$a/$ab/$dist/$dist-$version.pgz&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;readme&#34;</span><span class="p">:</span> <span class="s2">&#34;/dist/$a/$ab/$dist/$dist-$version.readme&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;meta&#34;</span><span class="p">:</span>   <span class="s2">&#34;/dist/$a/$ab/$dist/$dist-$version.json&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;by-dist&#34;</span><span class="p">:</span>      <span class="s2">&#34;/by/dist/$a/$ab/$dist.json&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;by-extension&#34;</span><span class="p">:</span> <span class="s2">&#34;/by/extension/$a/$ab/$extension.json&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;by-owner&#34;</span><span class="p">:</span>     <span class="s2">&#34;/by/owner/$a/$ab/$owner.json&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;by-manager&#34;</span><span class="p">:</span>   <span class="s2">&#34;/by/manager/$a/$ab/$manager.json&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>The PGXN client will always fetch this file before it does anything else,
because the file tells it how to find stuff. The advantage here is that the
client doesn&rsquo;t have to know anything about how the directory is actually
organized, just what the template variables might be. They are:</p>
<ul>
<li><code>$dist</code>: A <a href="https://github.com/theory/pgxn/wiki/PGXN-Meta-Spec#name">distribution name</a></li>
<li><code>$version</code>: A <a href="https://github.com/theory/pgxn/wiki/PGXN-Meta-Spec#version">version number</a></li>
<li><code>$extension</code>: An <a href="https://github.com/theory/pgxn/wiki/PGXN-Meta-Spec#provides">extension name</a></li>
<li><code>$owner</code>: An <a href="https://github.com/theory/pgxn/wiki/PGXN-Meta-Spec#owner">owner&rsquo;s name</a></li>
<li><code>$manager</code>: A release manager&rsquo;s name (managers are the people who upload
distributions to PGXN)</li>
<li><code>$a</code>: The first letter of a distribution, extension, owner, or manager name.</li>
<li><code>$ab</code>: The first two letters of a distribution, extension, owner, or manager
name.</li>
</ul>
<p>I&rsquo;m not thrilled about using prefix-staggering to avoid having too many files
in a directory. But the truth is that this approach allows me to punt. I could
also make sure the client supports, for example, <code>$bc</code> and <code>$cd</code>, so that one
could stagger things differently. And then the nice thing is that I don&rsquo;t have
to use those at all. The templates will tell the client exactly how to
construct the URIs for things, and the templates needn&rsquo;t include those
staggering variables if they&rsquo;re not appropriate. The client won&rsquo;t care because
it will have no built-in knowledge of how things are organized. It will have
to find out from <code>index.json</code>.</p>
<p>From the URI templates, you can now see where the other metadata will be
stored. For extension names, a hypothetical pgTAP distribution with two
extensions will have a JSON file for each extension:</p>
<pre tabindex="0"><code>/by/extension/p/pg/pgtap.json
/by/extension/s/sc/schematap.json
</code></pre><p>The <code>pgtap.json</code> file will look something like this:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="s2">&#34;stable&#34;</span><span class="err">:</span>   <span class="s2">&#34;0.25.0&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;testing&#34;</span><span class="err">:</span>  <span class="s2">&#34;0.26.0b1&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;unstable&#34;</span><span class="err">:</span> <span class="s2">&#34;0.30.0u&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;versions&#34;</span><span class="err">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.26.0b1&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;dist&#34;</span><span class="p">:</span> <span class="s2">&#34;pgtap&#34;</span><span class="p">,</span> <span class="nt">&#34;version&#34;</span><span class="p">:</span> <span class="s2">&#34;0.26.0b1&#34;</span><span class="p">,</span> <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="s2">&#34;testing&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.30.0u&#34;</span><span class="p">:</span>  <span class="p">{</span> <span class="nt">&#34;dist&#34;</span><span class="p">:</span> <span class="s2">&#34;pgtap&#34;</span><span class="p">,</span> <span class="nt">&#34;version&#34;</span><span class="p">:</span> <span class="s2">&#34;0.30.0u&#34;</span><span class="p">,</span>  <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="s2">&#34;unstable&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.25.0&#34;</span><span class="p">:</span>   <span class="p">{</span> <span class="nt">&#34;dist&#34;</span><span class="p">:</span> <span class="s2">&#34;pgtap&#34;</span><span class="p">,</span> <span class="nt">&#34;version&#34;</span><span class="p">:</span> <span class="s2">&#34;0.25.0&#34;</span><span class="p">,</span>   <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="s2">&#34;stable&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.24.0&#34;</span><span class="p">:</span>   <span class="p">{</span> <span class="nt">&#34;dist&#34;</span><span class="p">:</span> <span class="s2">&#34;pgtap&#34;</span><span class="p">,</span> <span class="nt">&#34;version&#34;</span><span class="p">:</span> <span class="s2">&#34;0.24.0&#34;</span><span class="p">,</span>   <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="s2">&#34;stable&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.25.0&#34;</span><span class="p">:</span>   <span class="p">{</span> <span class="nt">&#34;dist&#34;</span><span class="p">:</span> <span class="s2">&#34;pgtap&#34;</span><span class="p">,</span> <span class="nt">&#34;version&#34;</span><span class="p">:</span> <span class="s2">&#34;0.23.0&#34;</span><span class="p">,</span>   <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="s2">&#34;stable&#34;</span>  <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>Right at the top, it would always list the most recent stable, testing, and
unstable version number, and then it would have a list metadata for all
versions. Said metadata would include the associated distribution name,
version, and release status.</p>
<p>Here&rsquo;s how it would work. Say I ask the client to install pgtap:</p>
<pre tabindex="0"><code>PGXN&gt; install extension pgtap
</code></pre><p>The client would first fetch <code>/index.json</code>, then look for the URI template for
&ldquo;by-extension&rdquo;, which is <code>/by/extension/$a/$ab/$extension.json</code>. Filling in
the template, it would know to request <code>/by/extension/p/pg/pgtap.json</code>. With
that file, it would see that the most recent stable version is in the &ldquo;pgtap&rdquo;
distribution version 0.25.0. Using the <code>dist</code> URI template, which is
<code>/dist/$a/$ab/$dist-$version.pgz</code>, it would then fetch
<code>/dist/p/pg/pgtap/pgtap-0.25.0.pgz</code>.</p>
<p>The advantage here is that there are no symbolic links and no knowledge of the
directory structure built into clients. The client just knows to fetch
<code>/index.json</code> and then to use the templates in that file to fetch other
information. That&rsquo;s the whole interface. Very <a href="https://en.wikipedia.org/wiki/REST">REST</a>ful.</p>
<p>The structure of the other <code>/by</code> files would be similar. For</p>
<pre><code>PGXN&gt; install dist pgtap
</code></pre>
<p>the client would use the &ldquo;by-dist&rdquo; URI template to construct the URL
<code>/by/dist/p/pg/pgtap.json</code>. That file would have something like:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="s2">&#34;stable&#34;</span><span class="err">:</span>   <span class="s2">&#34;0.25.0&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;testing&#34;</span><span class="err">:</span>  <span class="s2">&#34;0.26.0b1&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;unstable&#34;</span><span class="err">:</span> <span class="s2">&#34;0.30.0u&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;versions&#34;</span><span class="err">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.26.0b1&#34;</span><span class="p">:</span> <span class="s2">&#34;testing&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.30.0u&#34;</span><span class="p">:</span>  <span class="s2">&#34;unstable&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.25.0&#34;</span><span class="p">:</span>   <span class="s2">&#34;stable&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.24.0&#34;</span><span class="p">:</span>   <span class="s2">&#34;stable&#34;</span> <span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;0.23.0&#34;</span><span class="p">:</span>   <span class="s2">&#34;stable&#34;</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>So then the client would know that &ldquo;0.25.0&rdquo; was the most recent version, and
use the <code>dist</code> URI template to request <code>/dist/p/pg/pgtap/pgtap-0.25.0.pgz</code>.</p>
<p>If The client command had been:</p>
<pre tabindex="0"><code>PGXN&gt; readme dist pgtap
</code></pre><p>It would use the <code>readme</code> URI template. And the command:</p>
<pre tabindex="0"><code>PGXN&gt; meta dist pgtap
</code></pre><p>Would use the <code>meta</code> URI template to fetch the metadata for the distribution.</p>
<p>If the client had requested a specific version:</p>
<pre tabindex="0"><code>PGXN&gt; install dist 0.23.0
</code></pre><p>It could either use the <code>by-dist</code> URI template to download the list of all
versions to see if 0.23.0 was valid, or just use the <code>dist</code> URI template to
try to download the distribution itself.</p>
<p>And finally, the owner and manager JSON files, such as</p>
<pre tabindex="0"><code>/owner/t/th/theory.json
</code></pre><p>Would look something like:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="s2">&#34;full_name&#34;</span><span class="err">:</span> <span class="s2">&#34;David Wheeler&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;email&#34;</span><span class="err">:</span> <span class="s2">&#34;theory@pgxn.org&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;uri&#34;</span><span class="err">:</span> <span class="s2">&#34;https://justatheory.com&#34;</span><span class="err">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;distributions&#34;</span><span class="err">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;pgtap&#34;</span><span class="p">:</span> <span class="p">[</span> <span class="s2">&#34;0.25.0&#34;</span><span class="p">,</span> <span class="s2">&#34;0.24.0&#34;</span><span class="p">,</span> <span class="s2">&#34;0.23.0&#34;</span> <span class="p">]</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;pair&#34;</span><span class="p">:</span> <span class="p">[</span> <span class="s2">&#34;0.2.0&#34;</span><span class="p">,</span> <span class="s2">&#34;0.1.0&#34;</span><span class="p">,</span> <span class="s2">&#34;0.0.5&#34;</span> <span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>With that, the client can be asked to fetch metadata for a given owner name
and use it to figure out what distributions and versions the the owner, um,
owns. One could then fetch the metadata, readme, or distribution file for any
of those distributions and versions.</p>
<p>Overall, I think that this is a much better solution than I outlined
<a href="https://blog.pgxn.org/post/954535657/thoughts-on-the-network-directory-structure">before</a>. If only I could figure out something more
elegant that the prefix-staggering/hashing stuff, it would be just about
perfect.</p>
<p>Thoughts?</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/954535657</id><title type="html">Thoughts on the Network Directory Structure</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2010/thoughts-on-the-network-directory-structure/"/><updated>2026-10-07T16:13:48Z</updated><published>2010-08-14T23:59:12Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="mirror" label="Mirror"/><category scheme="https://blog.pgxn.org/tags" term="directory" label="Directory"/><category scheme="https://blog.pgxn.org/tags" term="structure" label="Structure"/><category scheme="https://blog.pgxn.org/tags" term="metadata" label="Metadata"/><category scheme="https://blog.pgxn.org/tags" term="json" label="JSON"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="scalability" label="Scalability"/><category scheme="https://blog.pgxn.org/tags" term="rest" label="REST"/><summary type="html"><![CDATA[<p>I&rsquo;ve been thinking about the arrangement of stuff to be distributed to the
mirrors. In doing so, I&rsquo;ve kept three goals in mind:</p>
<ol>
<li>Allow things to scale.</li>
<li>Try to put things in intuitive locations.</li>
<li>Allow for metadata to be stored to act as a Web API.</li>
</ol>
<p>The first two are a bit mutually-contradicting. Thinking scalability mainly
means minimizing the chances that too many files can be put into a single
directory. <a href="https://www.cpan.org/misc/ZCAN.html" title="Zen of Comprehensive Archive Networks (look for “Naming”)">CPAN</a> started out with all author directories in a single
directory. Given the sheer number of CPAN contributors, this quickly got to be
a bottleneck. To correct for that, they started &ldquo;hashing&rdquo; the first two
letters of author names to create subdirectories. My CPAN directory, for
example, is <a href="https://www.cpan.org/authors/id/D/DW/DWHEELER/"><code>D/DW/DWHEELER</code></a>. That doesn&rsquo;t quite make for an intuitive
location, but it&rsquo;s not bad, and the tradeoff seems sufficient to keep things
sane.</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I&rsquo;ve been thinking about the arrangement of stuff to be distributed to the
mirrors. In doing so, I&rsquo;ve kept three goals in mind:</p>
<ol>
<li>Allow things to scale.</li>
<li>Try to put things in intuitive locations.</li>
<li>Allow for metadata to be stored to act as a Web API.</li>
</ol>
<p>The first two are a bit mutually-contradicting. Thinking scalability mainly
means minimizing the chances that too many files can be put into a single
directory. <a href="https://www.cpan.org/misc/ZCAN.html" title="Zen of Comprehensive Archive Networks (look for “Naming”)">CPAN</a> started out with all author directories in a single
directory. Given the sheer number of CPAN contributors, this quickly got to be
a bottleneck. To correct for that, they started &ldquo;hashing&rdquo; the first two
letters of author names to create subdirectories. My CPAN directory, for
example, is <a href="https://www.cpan.org/authors/id/D/DW/DWHEELER/"><code>D/DW/DWHEELER</code></a>. That doesn&rsquo;t quite make for an intuitive
location, but it&rsquo;s not bad, and the tradeoff seems sufficient to keep things
sane.</p>
<p>The storage of metadata is important, too, as I plan to have the <a href="https://pgxn.org/status.html#client">client</a> send
requests for JSON files to mirrors in order to find distributions. Basically,
this came down to a naming convention, as well as a recipe for how to find
metadata files for people and extensions. What I&rsquo;ve come up with is three
directories:</p>
<ul>
<li><code>meta</code> will contain PGXN metadata</li>
<li><code>dist</code> will contain distributions</li>
<li><code>by</code> will contain query-able JSON files</li>
</ul>
<p>The <code>meta</code> directory will contain PGXN metadata (<code>mirrors.json</code>, timestamp,
other detritus). If I ever create an index file that lists all distributions,
it would go there, too. I&rsquo;m not going to discuss it any further here, except
to note that it already has <code>mirrors.json</code>, which clients will be able to use
to present users with a list of mirrors to choose from.</p>
<p>The <code>dist</code> directory will be organized with directories named with &ldquo;hashes&rdquo; of
the first two letters of distribution names. Let&rsquo;s say that I&rsquo;m releasing
<a href="https://pgtan.org/">pgTAP</a> 0.25 on PGXN. To distribute it, the management application will create
the directory (if it doesn&rsquo;t already exist):</p>
<pre tabindex="0"><code>dist/p/pg/pgtap/
</code></pre><p>Then, for the 0.25 release, it will add three files to that directory:</p>
<pre tabindex="0"><code>dist/p/pg/pgtap/pgtap-0.25.pgz
dist/p/pg/pgtap/pgtap-0.25.json
dist/p/pg/pgtap/pgtap-0.25.readme
</code></pre><p>The <code>pgtap-0.25.pgz</code> file will contain the <a href="https://wiki.postgresql.org/wiki/PGXN#Distribution_Layout">zipped distribution</a> ready for
download. <code>pgtap-0.25.json</code> will contain metadata about the distribution, such
as the owner&rsquo;s name, the manager&rsquo;s name (more on these folks below), list of
included extensions, location of the <code>.pgz</code> and <code>.readme</code>, and its SHA1.
<code>pgtap-0.25.readme</code> will of course contain the <code>README</code> for the distribution
(if it has one).</p>
<p>Every release of pgTAP will have these three files, so after several releases,
the <code>pgtap</code> directory might have these files:</p>
<pre tabindex="0"><code>dist/p/pg/pgtap/pgtap-0.23.pgz
dist/p/pg/pgtap/pgtap-0.23.json
dist/p/pg/pgtap/pgtap-0.23.readme
dist/p/pg/pgtap/pgtap-0.24.pgz
dist/p/pg/pgtap/pgtap-0.24.json
dist/p/pg/pgtap/pgtap-0.24.readme
dist/p/pg/pgtap/pgtap-0.25.pgz
dist/p/pg/pgtap/pgtap-0.25.json
dist/p/pg/pgtap/pgtap-0.25.readme
dist/p/pg/pgtap/pgtap.json
</code></pre><p>The last file there, <code>pgtap.json</code>, will actually be a symlink to the JSON file
for latest production release of pgTAP. In this case, it would link to
<code>pgtap-0.25.json</code>. The nice thing about this is that, if a client wants to
find the information about the latest release of the pgtap distribution, all
it will have to do is send an HTTP GET request for
<code>dist/p/pg/pgtap/pgtap.json</code> to any mirror.</p>
<p>The <code>by</code> directory will also contain JSON files for clients to request. The
idea is that you want to find information &ldquo;by&rdquo; something. To start with, there
will be three subdirectories:</p>
<pre tabindex="0"><code>by/extension/
by/manager/
by/owner/
</code></pre><p>The first directory, <code>by/extension/</code>, will contain links to JSON files for
extensions. Say that the pgTAP distribution offers two extensions to
PostgreSQL named &ldquo;pgtap&rdquo; and &ldquo;schematap&rdquo;. The links would be:</p>
<pre tabindex="0"><code>by/extension/p/pg/pgtap/pgtap-0.23.json
by/extension/p/pg/pgtap/pgtap-0.24.json
by/extension/p/pg/pgtap/pgtap-0.25.json
by/extension/p/pg/pgtap/pgtap.json
</code></pre><p>Each of these will simply be symlinks pointing to the appropriate distribution
files:</p>
<pre tabindex="0"><code>dist/p/pg/pgtap/pgtap-0.23.json
dist/p/pg/pgtap/pgtap-0.24.json
dist/p/pg/pgtap/pgtap-0.25.json
dist/p/pg/pgtap/pgtap.json
</code></pre><p>Yes, the last one is a symlink to a symlink. Similarly, these files for
&ldquo;schematap&rdquo;:</p>
<pre tabindex="0"><code>by/extension/s/sc/schematap/schematap-0.23.json
by/extension/s/sc/schematap/schematap-0.24.json
by/extension/s/sc/schematap/schematap-0.25.json
by/extension/s/sc/schematap/schematap.json
</code></pre><p>Point to exactly the same files. Essentially, this is a way for extensions to
point to the distributions that contain them.
<code>by/extension/s/sc/schematap/schematap.json</code> points to
<code>dist/p/pg/pgtap/pgtap.json</code>, which contains path to the <code>.pgz</code> file to
download (and lots of other metadata, too).</p>
<p>The idea is that, whatever the name of the extension you want, the client will
be able to easily find the metadata file that tells it where to find the
distribution.</p>
<p>The <code>by/manager</code> and <code>by/owner</code> directories, on the other hand, contain JSON
files with information about managers and owners and their distributions.
Definitions:</p>
<p>An &ldquo;owner&rdquo; is someone who <em>owns</em> a distribution. This will often be the
original author of an extension, but may be someone else if maintenance has
been passed on. Basically, the &ldquo;owner&rdquo; is the person who should be contacted
with bug reports and the like</p>
<p>A &ldquo;manager&rdquo; is someone who manages the release process. This is the user who
will log into the PGXN management application and upload a distribution for
release.</p>
<p>These two people will often be the same person, but not always. I&rsquo;ve avoided
the term &ldquo;author&rdquo; (the term used by CPAN) because the author of an extension
may no longer maintain it. &ldquo;Owner&rdquo; seemed like a better choice (individual
distributions are free to describe their contributors however they wish).</p>
<p>As the <em>owner</em> of a few extensions on PGXN, I&rsquo;d have this file:</p>
<pre tabindex="0"><code>by/owner/t/th/theory.json
</code></pre><p>This file would contain a list of my distributions and perhaps some other
information (like my full name and blog URL).</p>
<p>As the <em>manager</em> of extensions on PGXN (that is, I actually uploaded them), I
would also have:</p>
<pre tabindex="0"><code>by/manager/t/th/theory.json
</code></pre><p>This file would contain a list of the distributions I&rsquo;ve released on PGXN.
This might be exactly the same as the list in my owner file, but may not be.
Perhaps for one release of pgTAP, say 0.26, <a href="https://leto.net/dukeleto.pl/">Duke Leto</a> uploaded a release. In
that case, 0.26 would probably be in my owner file, but not in my manager
file: it would be in Duke&rsquo;s manager file, instead.</p>
<p>So these are the basics of the directory structure for the networked mirrors.
Note that I haven&rsquo;t thought much about everything that will go into the JSON
files (or whether or not they&rsquo;d be versioned). That will likely depend quite a
lot on what the management database ends up looking like. I&rsquo;ll be working on
that next.</p>
<p>But other than that, comments? Questions? Criticisms? Recommendations? Leave a
comment and let me know!</p>
]]></content></entry></feed>