<?xml version="1.0" encoding="UTF-8" standalone="yes"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-us"><id>https://blog.pgxn.org/tags/indexing/</id><title>Indexing</title><updated>2011-03-22T05:13:00Z</updated><link rel="self" type="application/atom+xml" href="https://blog.pgxn.org/tags/indexing/feed.xml"/><link rel="alternate" type="text/html" href="https://blog.pgxn.org/tags/indexing/"/><author><name>The PGXN Maintainers</name></author><generator uri="https://gohugo.io/" version="0.167.0">Hugo</generator><entry><id>https://blog.pgxn.org/post/4018670551</id><title type="html">Thoughts on Indexing and Documentation</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/indexing-docs/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-03-22T05:13:00Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="indexing" label="Indexing"/><category scheme="https://blog.pgxn.org/tags" term="full-text-search" label="Full Text Search"/><category scheme="https://blog.pgxn.org/tags" term="readme" label="README"/><category scheme="https://blog.pgxn.org/tags" term="documentation" label="Documentation"/><category scheme="https://blog.pgxn.org/tags" term="search-results" label="Search Results"/><category scheme="https://blog.pgxn.org/tags" term="search" label="Search"/><summary type="html"><![CDATA[<p>So I&rsquo;m designing the full text indexing for the PGXN search site. I&rsquo;m modeling
it on <a href="https://http//search.cpan.org">CPAN Search</a>, which has been great. There are four search options:</p>
<ul>
<li>Full documentation search. This is the most common. Includes doc title and
body.</li>
<li>User search. Search on names, nicknames, email addresses, URIs, etc.</li>
<li>Distribution search. Search on distribution name, abstract, description,
tags, and the README.</li>
<li>Extension search. Search on extension name and abstract.</li>
<li>Tag search. Search on tag name only.</li>
</ul>
<p>The documentation search is the one I&rsquo;m perhaps least sure about. It assumes
that each extension in a distribution will have documentation. But so far that
has not really been the practice for PostgreSQL extensions. Most folks seem to
stick the documentation in the README. And even then it can be <a href="https://master.pgxn.org/dist/countnulls/1.0.0/README.txt">almost
nothing</a>. So a search for &ldquo;count nulls&rdquo; probably would not find &ldquo;countnulls&rdquo;
extension, because there is no documentation. What should I do about this? I&rsquo;m
thinking one of:</p>]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>So I&rsquo;m designing the full text indexing for the PGXN search site. I&rsquo;m modeling
it on <a href="https://http//search.cpan.org">CPAN Search</a>, which has been great. There are four search options:</p>
<ul>
<li>Full documentation search. This is the most common. Includes doc title and
body.</li>
<li>User search. Search on names, nicknames, email addresses, URIs, etc.</li>
<li>Distribution search. Search on distribution name, abstract, description,
tags, and the README.</li>
<li>Extension search. Search on extension name and abstract.</li>
<li>Tag search. Search on tag name only.</li>
</ul>
<p>The documentation search is the one I&rsquo;m perhaps least sure about. It assumes
that each extension in a distribution will have documentation. But so far that
has not really been the practice for PostgreSQL extensions. Most folks seem to
stick the documentation in the README. And even then it can be <a href="https://master.pgxn.org/dist/countnulls/1.0.0/README.txt">almost
nothing</a>. So a search for &ldquo;count nulls&rdquo; probably would not find &ldquo;countnulls&rdquo;
extension, because there is no documentation. What should I do about this? I&rsquo;m
thinking one of:</p>
<ul>
<li>
<p>Encourage folks to write documentation. I&rsquo;m going to do this anyway, because
the docs will really help the visibility of an extension on the site. It
looks <a href="https://theory.github.com/pgxn/pgtap.html">like this</a>. If you have no docs for an extension, your extension will
not appear in the search results (or perhaps it might, but link to the
distribution).</p>
</li>
<li>
<p>If there is no documentation for an extension in a distribution, index the
README as the documentation. I&rsquo;m not really keen on this idea, because the
README should describe the distribution, how to install it, etc. I&rsquo;m
planning to use it in the distribution-specific index. Documentation of the
extension should be more about how the extension works, what it&rsquo;s interface
is, etc. Or so it seems to me, at least (I&rsquo;m admittedly biased to this
practice among CPAN modules). But at least with this approach there would be
a link to &ldquo;documentation&rdquo; for an extension on the search site.</p>
</li>
</ul>
<p>Erm, not really thinking of any other options. I feel pretty strongly that
folks should write docs for their extensions, as much as possible, and I&rsquo;ve
set things up so that, from PGXN&rsquo;s point of view, at least, you can write
documentation in whatever format you like (assuming the format is supported by
or added to <a href="https://search.cpan.org/perldoc?Text::Markup">Text::Markup</a>), as long as they&rsquo;re in a <code>doc/</code> or <code>docs</code>
directory. I want it to be as easy as possible. But I also want there to be
decent search results ASAP.</p>
<p>Comments?</p>
]]></content></entry><entry><id>https://blog.pgxn.org/post/3641963548</id><title type="html">Question: Ajax and Search Engine Indexing</title><link rel="alternate" type="text/html" href="https://blog.pgxn.org/2011/ajax-search-indexing/"/><updated>2026-10-07T16:13:48Z</updated><published>2011-03-04T20:20:08Z</published><author><name>David E. Wheeler</name></author><category scheme="https://blog.pgxn.org/tags" term="ajax" label="Ajax"/><category scheme="https://blog.pgxn.org/tags" term="api" label="API"/><category scheme="https://blog.pgxn.org/tags" term="search-engine" label="Search Engine"/><category scheme="https://blog.pgxn.org/tags" term="indexing" label="Indexing"/><category scheme="https://blog.pgxn.org/tags" term="indexability" label="Indexability"/><category scheme="https://blog.pgxn.org/tags" term="response-code" label="Response Code"/><category scheme="https://blog.pgxn.org/tags" term="not-found" label="Not Found"/><summary type="html"><![CDATA[I&rsquo;ve started working on the main (search) site in earnest now. The basic
layout is done, and I&rsquo;m working on the distribution view (<a href="https://theory.github.com/pgxn/pgtapdist.html">mockup</a>). My
thinking so far has been that I would simply serve a page that requested, say,
<code>/dist/pgTAP/</code>, and that page would use Ajax requests to fetch the data from
the <a href="https://api.pgxn.org/">API server</a> and display stuff. I think this will work pretty well except
for one thing: 404s.]]></summary><content type="html" xml:base="https://blog.pgxn.org/" xml:space="preserve"><![CDATA[<p>I&rsquo;ve started working on the main (search) site in earnest now. The basic
layout is done, and I&rsquo;m working on the distribution view (<a href="https://theory.github.com/pgxn/pgtapdist.html">mockup</a>). My
thinking so far has been that I would simply serve a page that requested, say,
<code>/dist/pgTAP/</code>, and that page would use Ajax requests to fetch the data from
the <a href="https://api.pgxn.org/">API server</a> and display stuff. I think this will work pretty well except
for one thing: 404s.</p>
<p>That is, if you request <code>/dist/nonexistent/</code>, then it will load a page with
the HTTP status code <code>200 OK</code>, but then, when the Ajax request 404s, it will
show a &ldquo;Not found&rdquo; error message. That&rsquo;s all well and good, but I&rsquo;m wondering
about the impact of two things:</p>
<ol>
<li>
<p>Since the page itself won&rsquo;t 404, search engines might index links to
nonexistent extensions. Of course, bad links won&rsquo;t be <em>that</em> common, but
of course they do happen and then tend to live forever.</p>
</li>
<li>
<p>If the search site uses Ajax to fetch the contents of a page via JSON (or,
for documentation, as an HTML document it will put into a div), will the
full content be properly indexed by search engines?</p>
</li>
</ol>
<p>So these are serious questions, in my mind. Do we loose good search engine
indelibility when we load content dynamically?</p>
<p>Of course, I can instead write it so that the back end fetches stuff from the
API server (and perhaps directly from the file system) and get &lsquo;round these
issues, but then it&rsquo;s less of a cool example of the use of the API server.</p>
<p>What do you think? Good advice much appreciated!</p>
]]></content></entry></feed>