<?xml version="1.0" encoding="UTF-8" standalone="yes"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-us"><id>https://blog.pgxn.org/tags/search/</id><title>Search</title><updated>2011-03-22T05:13:00Z</updated><link rel="self" type="application/atom+xml" href="https://blog.pgxn.org/tags/search/feed.xml"/><link rel="alternate" type="text/html" href="https://blog.pgxn.org/tags/search/"/><author><name>The PGXN Maintainers</name></author><generator uri="https://gohugo.io/" version="0.167.0">Hugo</generator><entry><id>https://blog.pgxn.org/post/4018670551</id><title type="html">Thoughts on Indexing and Documentation</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/indexing-docs/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-03-22T05:13:00Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="indexing" label="Indexing"/><category scheme="https://blog.pgxn.org/tags" term="full-text-search" label="Full Text Search"/><category scheme="https://blog.pgxn.org/tags" term="readme" label="README"/><category scheme="https://blog.pgxn.org/tags" term="documentation" label="Documentation"/><category scheme="https://blog.pgxn.org/tags" term="search-results" label="Search Results"/><category scheme="https://blog.pgxn.org/tags" term="search" label="Search"/><summary type="html"><![CDATA[<p>So I&rsquo;m designing the full text indexing for the PGXN search site. I&rsquo;m modeling
it on <a href="https://http//search.cpan.org">CPAN Search</a>, which has been great. There are four search options:</p>
<ul>
<li>Full documentation search. This is the most common. Includes doc title and
body.</li>
<li>User search. Search on names, nicknames, email addresses, URIs, etc.</li>
<li>Distribution search. Search on distribution name, abstract, description,
tags, and the README.</li>
<li>Extension search. Search on extension name and abstract.</li>
<li>Tag search. Search on tag name only.</li>
</ul>
<p>The documentation search is the one I&rsquo;m perhaps least sure about. It assumes
that each extension in a distribution will have documentation. But so far that
has not really been the practice for PostgreSQL extensions. Most folks seem to
stick the documentation in the README. And even then it can be <a href="https://master.pgxn.org/dist/countnulls/1.0.0/README.txt">almost
nothing</a>. So a search for &ldquo;count nulls&rdquo; probably would not find &ldquo;countnulls&rdquo;
extension, because there is no documentation. What should I do about this? I&rsquo;m
thinking one of:</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>So I&rsquo;m designing the full text indexing for the PGXN search site. I&rsquo;m modeling
it on <a href="https://http//search.cpan.org">CPAN Search</a>, which has been great. There are four search options:</p>
<ul>
<li>Full documentation search. This is the most common. Includes doc title and
body.</li>
<li>User search. Search on names, nicknames, email addresses, URIs, etc.</li>
<li>Distribution search. Search on distribution name, abstract, description,
tags, and the README.</li>
<li>Extension search. Search on extension name and abstract.</li>
<li>Tag search. Search on tag name only.</li>
</ul>
<p>The documentation search is the one I&rsquo;m perhaps least sure about. It assumes
that each extension in a distribution will have documentation. But so far that
has not really been the practice for PostgreSQL extensions. Most folks seem to
stick the documentation in the README. And even then it can be <a href="https://master.pgxn.org/dist/countnulls/1.0.0/README.txt">almost
nothing</a>. So a search for &ldquo;count nulls&rdquo; probably would not find &ldquo;countnulls&rdquo;
extension, because there is no documentation. What should I do about this? I&rsquo;m
thinking one of:</p>
<ul>
<li>
<p>Encourage folks to write documentation. I&rsquo;m going to do this anyway, because
the docs will really help the visibility of an extension on the site. It
looks <a href="https://theory.github.com/pgxn/pgtap.html">like this</a>. If you have no docs for an extension, your extension will
not appear in the search results (or perhaps it might, but link to the
distribution).</p>
</li>
<li>
<p>If there is no documentation for an extension in a distribution, index the
README as the documentation. I&rsquo;m not really keen on this idea, because the
README should describe the distribution, how to install it, etc. I&rsquo;m
planning to use it in the distribution-specific index. Documentation of the
extension should be more about how the extension works, what it&rsquo;s interface
is, etc. Or so it seems to me, at least (I&rsquo;m admittedly biased to this
practice among CPAN modules). But at least with this approach there would be
a link to &ldquo;documentation&rdquo; for an extension on the search site.</p>
</li>
</ul>
<p>Erm, not really thinking of any other options. I feel pretty strongly that
folks should write docs for their extensions, as much as possible, and I&rsquo;ve
set things up so that, from PGXN&rsquo;s point of view, at least, you can write
documentation in whatever format you like (assuming the format is supported by
or added to <a href="https://search.cpan.org/perldoc?Text::Markup">Text::Markup</a>), as long as they&rsquo;re in a <code>doc/</code> or <code>docs</code>
directory. I want it to be as easy as possible. But I also want there to be
decent search results ASAP.</p>
<p>Comments?</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/3099288750</id><title type="html">PGXN API RFC</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/pgxn-api-rfc/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-02-04T04:05:33Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="search" label="Search"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="mirror" label="Mirror"/><category scheme="https://blog.pgxn.org/tags" term="json" label="JSON"/><category scheme="https://blog.pgxn.org/tags" term="cpan" label="CPAN"/><category scheme="https://blog.pgxn.org/tags" term="metacpan" label="metacpan"/><category scheme="https://blog.pgxn.org/tags" term="javascript" label="JavaScript"/><category scheme="https://blog.pgxn.org/tags" term="application" label="Application"/><summary type="html"><![CDATA[<p>Things slowed up a bit over the last couple of months, I admit. There are any
number of reasons for that, not the least were the intrusion of the holidays
and a <a href="https://www.designsceneapp.com/">little side project</a> I&rsquo;ve been hacking on after-hours (and sometimes
during-hours). But I&rsquo;m ramping things up again now and need <em>your</em> feedback on
my current plans. Here&rsquo;s what I&rsquo;m working on: the search site.</p>
<h2 id="search-sites-and-apis">Search Sites and APIs</h2>
<p>Well, sort of. First of all, I&rsquo;ve decided that the &ldquo;search site&rdquo; should not be
a separate thing. The <a href="https://pgxn.org/">main site</a> will be the search site. This is following
the example of <a href="https://openjsan.org">JSAN</a>, as well as feedback from <a href="https://search.cpan.org/~gbarr/">Graham Barr</a>, who created and
maintains <a href="https://search.cpan.org/">CPAN Search</a>. Apparently people are often confused that
<a href="https://search.cpan.org/">search.cpan.org</a> is separate from <a href="https://www.cpan.org">www.cpan.org</a>. No point in
adding in confusion from the beginning. And besides, now that the PGXN
fund-raising <a href="https://blog.pgxn.org/post/2063677299/goooooooaaaal">is over</a>, I don&rsquo;t know what else would go on the home page.</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>Things slowed up a bit over the last couple of months, I admit. There are any
number of reasons for that, not the least were the intrusion of the holidays
and a <a href="https://www.designsceneapp.com/">little side project</a> I&rsquo;ve been hacking on after-hours (and sometimes
during-hours). But I&rsquo;m ramping things up again now and need <em>your</em> feedback on
my current plans. Here&rsquo;s what I&rsquo;m working on: the search site.</p>
<h2 id="search-sites-and-apis">Search Sites and APIs</h2>
<p>Well, sort of. First of all, I&rsquo;ve decided that the &ldquo;search site&rdquo; should not be
a separate thing. The <a href="https://pgxn.org/">main site</a> will be the search site. This is following
the example of <a href="https://openjsan.org">JSAN</a>, as well as feedback from <a href="https://search.cpan.org/~gbarr/">Graham Barr</a>, who created and
maintains <a href="https://search.cpan.org/">CPAN Search</a>. Apparently people are often confused that
<a href="https://search.cpan.org/">search.cpan.org</a> is separate from <a href="https://www.cpan.org">www.cpan.org</a>. No point in
adding in confusion from the beginning. And besides, now that the PGXN
fund-raising <a href="https://blog.pgxn.org/post/2063677299/goooooooaaaal">is over</a>, I don&rsquo;t know what else would go on the home page.</p>
<p>The other thing that&rsquo;s happened is, just as I was getting my butt in gear on
this stuff, a new CPAN search site came to my attention, <a href="https://search.metacpan.org/">μετα CPAN</a>. This is
an interesting project. What they did instead of creating a monolithic HTML
search site is to create a <a href="https://github.com/CPAN-API/cpan-api/wiki/API-docs">simple API</a> that serves nothing but JSON. It has
search and displays metadata for CPAN objects (distributions, maintainers,
modules, etc.). The search site, then, is not really a site at all, but a pure
JavaScript application. Once you load it, it just uses the API server to get
all the data. There are a few tricks server-side to proxy the API server so as
to avoid cross-site scripting issues. But otherwise it just works in the
browser.</p>
<p>Now I&rsquo;m not sure I&rsquo;ll do the same thing, exactly, but there&rsquo;s a lot of appeal
in creating a RESTful API server that&rsquo;s independent of the search site, and
then building the search site to use it. It also has the advantage of being
useful for other projects to just use. Want to create a PGXN search widget for
your blog? Yeah, there&rsquo;s an API for that.</p>
<h2 id="a-super-restful-directory">A Super RESTful Directory</h2>
<p>Of course, thanks to the &ldquo;RESTful Directory&rdquo; design for the mirrors (described
<a href="https://blog.pgxn.org/post/988613682/restful-directory" title="A RESTful Directory">here</a> and revised <a href="https://blog.pgxn.org/post/1138292188/arch-and-extension-json" title="Architecture, New Extension JSON Format">here</a>), any mirror is a lightweight API already.
There&rsquo;s a <em>lot</em> of metadata one can get just from the static JSON files it
generates. The design is flexible&ndash;but designed with a command-line client in
mind. As such, many commands executed in a command-line client would likely
requires multiple requests to a mirror. For example:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-console" data-lang="console"><span class="line"><span class="cl"><span class="gp">&gt;</span> install pgtap
</span></span></code></pre></div><p>This would request <code>/by/extension/pgtap.json</code> from the server. It would then
parse that file and see that the latest stable version of pgTAP is in the
distribution &ldquo;pgTAP&rdquo; at version &ldquo;0.25.0&rdquo;. So it would then download
<code>/dist/pgTAP-0.25.0.pgz</code> to install.</p>
<p>This is great for a command-line client, but wouldn&rsquo;t be so great for a search
site to be responsive. Ideally, a site should send a single request to get all
the data it needs for a particular page.</p>
<p>So here&rsquo;s what I&rsquo;m thinking for a PGXN API server: It will offer a superset of
the functionality of any other PGXN mirror. That is, all the JSON files in a
mirror will be present, but many of them will have more information than they
would on other mirrors. And then, of course, there will be other URIs to offer
additional API calls.</p>
<h3 id="details">Details</h3>
<p>So what does that look like? Let&rsquo;s take the pgTAP distribution, which I
released on PGXN earlier this week. To find the pgTAP distribution, one
requests:</p>
  <blockquote>
    <p><a href="https://master.pgxn.org/by/dist/pgTAP.json"><code>/by/dist/pgTAP.json</code></a></p>
  </blockquote>
<p>From that, one can see that the latest table release is 0.25.0, and so one can
then request</p>
  <blockquote>
    <p><a href="https://master.pgxn.org/dist/pgTAP/pgTAP-0.25.0.json"><code>/dist/pgTAP/pgTAP-0.25.0.json</code></a></p>
  </blockquote>
<p>to get all the metadata for that particular release. What I propose, to avoid
the two requests, is to include the contents of the second file in the first.
That would then have all the data necessary to generate the <a href="https://theory.github.com/pgxn/pgtapdist.html">pgTAP
distribution page</a> on the PGXN site.</p>
<p>The API would offer similar supersets of data for the <a href="https://master.pgxn.org/by/extension/pgtap.json">extension</a> , <a href="https://master.pgxn.org/by/owner/theory.json">owner</a> ,
and <a href="https://master.pgxn.org/by/tag/testing.json">tag</a> metadata files, to have the data necessary for the design of the
corresponding <a href="https://theory.github.com/pgxn/pgtap.html">extension</a>, <a href="https://theory.github.com/pgxn/theory.html">owner</a> and tag layouts of the site.</p>
<h3 id="additional-resources">Additional Resources</h3>
<p>In addition to adding metadata to the existing mirrored JSON files, there
would be other resources available for request from the API server. They would
include:</p>
<ul>
<li>
<p>Extension Documentation. Each distribution may include documentation for
included extensions in the <code>doc</code> subdirectory. These will go under the
directory for a specific distribution such as
<code>/dist/pgTAP/pgTAP-0.35.0/doc/pgtap.html</code>. The latest version of each
document would also be available under <code>/by/extension</code>, as in
<code>/by/extension/pgtap.html</code>. This requires that the documentation file have
the same base name as the extension file itself.</p>
</li>
<li>
<p>Other documentation. I&rsquo;d like to support arbitrary documentation, such as
for included binary executables, HOWTOs, etc. The canonical copies will go
under the versioned distribution URL, of course, but I&rsquo;m not sure about
permalinks. That might require an extension of the <a href="https://pgxn.org/meta/spec.txt">Meta Spec</a>; I haven&rsquo;t
quite figured that out, yet.</p>
</li>
<li>
<p>Source code. There will be an interface to browse an unpacked copy of any
distribution as plain text. This will be under <code>/src</code>, as in
<code>/src/pgTAP/pgTAP-0.35.0/</code>.</p>
</li>
</ul>
<h3 id="search-api">Search API</h3>
<p>Of course. This is the big one, really. I think it makes sense to have the
<code>/by</code> URI respond to search requests. Thus, a request for</p>
<pre tabindex="0"><code>/by?q=testing
</code></pre><p>would search everything. If you only want to search a certain category of
object, you&rsquo;d hit the appropriate URI:</p>
<pre tabindex="0"><code>/by/dist?q=tap
/by/owner?q=clochard
/by/tag?q=test
/by/extension?q=gis
</code></pre><p>The nice thing about this is that it retains the existing entity URLs. The
directory level determines which entities you get.</p>
<h2 id="your-thoughts">Your Thoughts?</h2>
<p>So that&rsquo;s my thinking on the search API. I&rsquo;m going to start hacking on it in
earnest tomorrow, and perhaps next week I can get a very early version out
(basically just another mirror to start with).</p>
<p>But what do you think? Seem like a sane approach? Am I missing anything
obvious or doing anything clearly stupid? Please let me know in the comments!</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/1473157261</id><title type="html">PGWest, Search Site Sneak Peak</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2010/pgwest-and-sneak-peak/"/><updated>2026-10-07T16:13:48Z</updated><published>2010-11-03T21:09:00Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="search" label="Search"/><category scheme="https://blog.pgxn.org/tags" term="pgwest" label="PgWest"/><category scheme="https://blog.pgxn.org/tags" term="presentation" label="Presentation"/><category scheme="https://blog.pgxn.org/tags" term="extensions" label="Extensions"/><category scheme="https://blog.pgxn.org/tags" term="postgresql-9.1" label="Postgresql 9.1"/><summary type="html"><![CDATA[I&rsquo;ve been working on my <a href="https://www.postgresqlconference.org/content/building-and-distributing-postgresql-extensions-without-learning-c">PostgreSQL Conference West presentation</a>, which
heavily features PGXN, of course. I think it&rsquo;s looking good. If you&rsquo;re at
<a href="https://www.postgresqlconference.org/2010/west/">PGWest</a> or are in the San Francisco area and free, come see the talk! Should
be a good introduction to creating PostgreSQL extensions and distributing them
on PGXN. The latest bit I added is a section on the modifications needed to
support the forthcoming <a href="https://s.coop/pgext"><code>CREATE EXTENSION</code></a> support slated for 9.1.
Fortunately it&rsquo;s dead simple, and will make dealing with extensions in the
database a lot simpler, administratively. Really looking forward to that. Of
course I&rsquo;ll post slides once the talk is over.]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I&rsquo;ve been working on my <a href="https://www.postgresqlconference.org/content/building-and-distributing-postgresql-extensions-without-learning-c">PostgreSQL Conference West presentation</a>, which
heavily features PGXN, of course. I think it&rsquo;s looking good. If you&rsquo;re at
<a href="https://www.postgresqlconference.org/2010/west/">PGWest</a> or are in the San Francisco area and free, come see the talk! Should
be a good introduction to creating PostgreSQL extensions and distributing them
on PGXN. The latest bit I added is a section on the modifications needed to
support the forthcoming <a href="https://s.coop/pgext"><code>CREATE EXTENSION</code></a> support slated for 9.1.
Fortunately it&rsquo;s dead simple, and will make dealing with extensions in the
database a lot simpler, administratively. Really looking forward to that. Of
course I&rsquo;ll post slides once the talk is over.</p>
<p>As part of preparing for the talk, and because there isn&rsquo;t currently much to
actually <em>see</em> of PGXN, I&rsquo;ve been mocking up the layout for the new search
site, which as you know from the <a href="https://pgxn.org/status.html">status page</a> is the next part of the project
I&rsquo;m slated to work on. I&rsquo;ve been committing the mockps to the gh-pages branch
of the repository, which means you can see what it looks like live on the net
<a href="https://theory.github.com/pgxn/">right here</a>. That&rsquo;s the home page, including our sponsor links and tag cloud.
Click the &ldquo;PGXN Search&rdquo; button to see a mockup of search results (or get them
<a href="https://theory.github.com/pgxn/results.html">here</a>). Click on any search result to see the mockup of a documentation page
(or link it <a href="https://theory.github.com/pgxn/pgtap.html">here</a>). The design is based on the <a href="https://www.oswd.org/design/preview/id/2839">lazydays</a> open-source Web
design, and I&rsquo;m quite happy with it. Your thoughts?</p>
<p>As this comes together, I&rsquo;m gearing up to start hacking on the app to produce
the search site. At this point, I&rsquo;m thinking that it would become the new home
page for <a href="https://pgxn.org/">PGXN</a>, rather than a separate search.pgxn.org site. Thoughts?</p>
<p>I&rsquo;ll post the slides tomorrow.</p>
]]></content></entry></feed>