Skip to content

Three new Halloween themes. See them

Portage on other sites

What our programs read, and how to say no.

Portage uses six small programs that visit other websites for the pages people make here. This page explains what each one does and how to block it if you’d rather it stayed away, and what we ask of the AI crawlers that visit us.

Last changed 7 October 2026

The short version

Portage sends six programs to other sites: PortageImport, PortagePreview, PortageLinkCheck, PortageFeed, PortageBetaCheck and PortageBot. Each one names itself in its user agent and links back to this page. They don’t sign in, keep cookies, fill in forms or get past bot checks, and they only read what anyone can see in a browser.

They keep their visits small and infrequent, and they stop as soon as a site says no.

One exception, said plainly: when you ask, PortageBot reads your own public Instagram or TikTok profile, and the link page it points to, once, even where those sites ask automated readers not to. It keeps only what goes on your page. Its own section below says exactly what it does.

On this page12 sections
  1. The short version
  2. PortageImport: moving an existing link page over
  3. Blocking PortageImport with robots.txt
  4. PortagePreview: a picture for a link
  5. Blocking PortagePreview with robots.txt
  6. PortageLinkCheck: broken links, and links back
  7. PortageFeed: the feeds a page shows
  8. PortageBetaCheck: whether a TestFlight beta has room
  9. Blocking PortageBetaCheck with robots.txt
  10. PortageBot: your own Instagram or TikTok, when you ask
  11. Asking us to stop
  12. What we ask of AI crawlers on Portage

PortageImport: moving an existing link page over

User agent: PortageImport/1.0 (+https://portage.camp/fetcher; a one-time read someone asked for).

What it reads: the one page someone pastes in, so we can rebuild it on Portage. That’s either a signed-in Portage user who says the page is theirs, or a visitor to our home page or sign-up page who wants to see it rebuilt before signing up. It reads the site’s robots.txt first. After that, it copies the profile photo that page shows, and for a signed-in user up to 16 of the small pictures beside its links as well, four at a time, only when the page itself came back within five seconds and never past eight seconds. Each is saved as a new file of our own, so nothing of the original file is kept. For a signed-in user they go into our storage, and any they don’t keep are deleted within two days. A visitor’s preview copies the profile photo and at most six of the first pictures on the page (the small pictures beside its links, and a video’s or a song’s still, which it asks YouTube, Vimeo, Spotify or SoundCloud for), all inside the same few seconds. It keeps them with the preview for thirty minutes, and puts only the photo in our storage, if they sign up and open the preview in their editor.

Starting from a profile: someone with no link page can paste one of their own profiles instead. For a GitHub, YouTube, Bandcamp or Mastodon profile, PortageImport reads that one profile page, robots.txt first, and keeps only the links the person listed there as their own. If one of those is a link page, it reads that page too, the same way. An Instagram or TikTok profile is read by PortageBot instead, described below. It never checks whether a name exists on any other site: suggestions like the same name on TikTok are only offered, for the person to confirm.

A page made for someone: when Portage makes a page for a person to claim, PortageImport may copy one photo of them from a website that allows it, such as their shop’s own team page. It reads that site’s robots.txt first and stays away if it says no.

When: only when someone asks, by pressing the import button in the editor, or by pasting an address on our home page or sign-up page. It doesn’t run on a schedule or come back later to check for changes.

How often: one visit to the page each time someone asks. One person can import up to 12 times an hour and 40 times a day. One network address can ask for 10 previews an hour, all previews together stop at 1,500 a day, and a page previewed in the last ten minutes isn’t visited again. All of Portage together, imports, previews and PortagePreview alike, makes at most 600 visits an hour to any one site. If a site says no, it doesn’t try again.

What it keeps: nothing from the page itself. The page is thrown away once it’s been read, and we don’t log its address. Only what the person decides to add ends up on their page. A preview keeps what was found in the page (the name, the links and the address they came from) for thirty minutes, so signing up doesn’t need a second visit, and after that it can’t be opened.

It follows up to three redirects, only over https, and checks robots.txt again before each one, on whatever site the redirect leads to. It reads at most 1.5 MB. If a page shows a bot check meant for people, it leaves it alone.

Blocking PortageImport with robots.txt

PortageImport reads robots.txt the standard way (RFC 9309), from the root of your site. It looks for rules for PortageImport first, then for rules that apply to every bot. The longest matching rule wins, and if an Allow and a Disallow tie, Allow wins.

If a site has no robots.txt, PortageImport goes ahead. If the robots.txt can’t be reached or returns a server error, it stays away.

We keep a copy of your robots.txt for a day, so changes take effect within 24 hours.

To block it from only part of your site, such as your profile pages, disallow just that path.

To block it everywhere, add these two lines to your robots.txt:

User-agent: PortageImport
Disallow: /

PortagePreview: a picture for a link

User agent: PortagePreview/1.0 (+https://portage.camp/fetcher; one read someone asked for).

What it reads: one page, when someone editing a Portage page presses “Find a picture from this link” beside one of their own links. It reads that site’s robots.txt first, then only the top of the page, at most 512 KB and stopping where the page’s head ends, looking for the picture and title the page names for link previews (its Open Graph or Twitter card tags). If the page names a picture, it reads the robots.txt of the site the picture is on, then copies that one picture, at most 8 MB, and saves it as a new file of our own, so nothing of the original file is kept.

The person then decides whether to put it on their link. Nothing changes on their page until they do, and a copy they don’t use is deleted within two days.

It never looks at a link in an 18+ section, a link to an 18+ site, or a picture on one.

When: only when someone presses that button. It doesn’t run when a link is typed or saved, it doesn’t run on a schedule, and it never comes back later to check for changes.

How often: one visit to the page and one to its picture each time someone asks. One person can ask 40 times an hour and 200 times a day. It counts towards the same limits as PortageImport, so all of Portage together still makes at most 600 visits an hour to any one site, and copies at most 600 pictures an hour from any one site.

What it keeps: nothing from the page itself. The page is thrown away once it’s been read, and we don’t log its address. Only the copy of the picture is kept, and the page’s title if the person chooses to use it for their link.

It follows up to three redirects, only over https, and checks robots.txt again before each one, on whatever site the redirect leads to. If a page shows a bot check meant for people, it leaves it alone.

Blocking PortagePreview with robots.txt

PortagePreview reads robots.txt the same way PortageImport does: rules for PortagePreview first, then rules that apply to every bot, with the longest matching rule winning. If a site has no robots.txt it goes ahead, and if the robots.txt can’t be reached it stays away. We keep a copy of your robots.txt for a day.

A site that blocks it never has a page read or a picture copied for a preview, and a link to it on a Portage page works exactly as before.

To block it everywhere, add these two lines to your robots.txt:

User-agent: PortagePreview
Disallow: /

PortageLinkCheck: broken links, and links back

User agent: PortageLinkCheck/1.0 (+https://portage.camp/fetcher).

What it checks: every link on every published Portage page, once a week, to see if it still works. It asks for the headers only, and only asks for the full page if that first answer says the page is gone, because some servers get header requests wrong. We only tell the person who made the page that a link is broken after it returns a 404 or 410 twice, a week apart, and we never remove or change the link.

It also checks the public profiles listed on a Portage page once a week, to see if they link back to it. That’s how the free verified mark works. It reads at most 512 KB of each profile and only looks at the links.

It doesn’t read robots.txt, because it only does what a visitor does when they tap a link: see whether the address still works, or whether a public profile links back.

To block it, answer its user agent with a 403. It treats that as “couldn’t check”, so nothing about your site gets reported and the page linking to you stays as it is.

PortageFeed: the feeds a page shows

User agent: PortageFeed/1.0 (+https://portage.camp/fetcher).

What it reads: feeds and other public addresses that someone added to their Portage page, so the page stays up to date. That covers the newest podcast episode, recent posts from a blog or newsletter, dates from a public calendar, release notes from an app’s feed and a post’s public preview, which is fetched when the page is saved. If someone pastes in a picture, it’s copied once when they save the page.

When and how often: whenever someone opens the page, but every visitor shares the same copy, which we keep for 15 minutes for a feed and an hour for a calendar. So we read any one feed at most four times an hour, however many people visit.

Like most feed readers, it doesn’t check robots.txt. A feed is published so programs can read it, and it’s only read because someone added it to their page.

To block it, answer its user agent with a 403. That part of the page then looks the same as it does when a feed is down.

PortageBetaCheck: whether a TestFlight beta has room

User agent: PortageBetaCheck/1.0 (+https://portage.camp/fetcher).

What it reads: one site, testflight.apple.com. When an app card on a Portage page has a TestFlight public link, it reads that link’s join page to see whether the beta is open, full or closed, so the card can say “The beta is full” instead of offering a button that won’t let anyone in. It reads robots.txt first, then the join page, at most 300 KB, and looks only at the one sentence that says whether the beta has room.

When and how often: once a day in our nightly run, and when the person who made the page opens that app card in their editor and the link hasn’t been read that day. However many pages share a link, it’s read at most once a day, one link at a time, with a pause between each.

What it keeps: the code at the end of the link, one word for what the page said (open, full, closed or gone) and when it was read. Nothing about the app, its developer or its testers. A link that isn’t on any page is forgotten after a month.

It doesn’t sign in, keep cookies or run the page’s script, and it only follows a redirect that stays on testflight.apple.com.

Blocking PortageBetaCheck with robots.txt

PortageBetaCheck reads robots.txt the same way PortageImport does: rules for PortageBetaCheck first, then rules that apply to every bot, with the longest matching rule winning. If there’s no robots.txt it goes ahead, and if the robots.txt can’t be reached it stays away until it can.

While it’s blocked, nothing is read, and within three days every app card is back to its usual TestFlight button.

To block it everywhere, add these two lines to your robots.txt:

User-agent: PortageBetaCheck
Disallow: /

PortageBot: your own Instagram or TikTok, when you ask

User agent: Mozilla/5.0 (compatible; PortageBot/1.0; +https://portage.camp/fetcher).

What it reads: when someone types their own Instagram or TikTok username on our home page or sign-up page, or pastes their own profile there, it reads that one public profile page, signed out, once: the name, the bio, the profile photo and the link or links in the bio. If one of those links is a link page (Linktree, Beacons, Bento, Carrd, Komi and the like), it reads that page once too, with PortageImport, the same way PortageImport reads any page someone brings over.

Honestly: Instagram and TikTok ask automated readers not to read their pages, and so do some link-page sites. PortageBot reads anyway, only in this one case: one person asking for their own profile, once, so they can start their page here without copying everything across by hand. Everywhere else, all our programs obey robots.txt exactly as this page says.

It never signs in, never uses anyone’s account, cookies or password, never goes through a proxy, and never solves or gets round a check meant for people. If a profile asks it to sign in, or answers with a check or a refusal, it stops there and says so, and the page starts from the username alone, with the way to paste a link page open.

When: only when someone presses the button. It doesn’t run on a schedule and never comes back to check for changes.

How often: one visit per profile each time someone asks. One network address can ask 10 times an hour, all of Portage together stops at 1,000 a day, and a profile read in the last ten minutes isn’t visited again.

What it keeps: nothing except what goes on that person’s page. The profile page is thrown away once it’s been read. The photo is copied and saved as a new file of our own. What was found is kept with the preview for thirty minutes, so signing up doesn’t need a second visit, and after that it can’t be opened.

To stop it reading your site, email us (below). Instagram and TikTok can turn it away by refusing it, which it respects at once.

Asking us to stop

Email us with your site’s address and which program you’d like us to stop. We keep a list of sites PortageImport and PortagePreview won’t read, and adding yours takes effect with our next update, usually within a day. You don’t need to tell us why.

If one of them visits more often or reads more than this page says, that’s a bug, and we’d like to hear about it.

What we ask of AI crawlers on Portage

It works the other way too. The pages people make on Portage are theirs, so our robots.txt asks the crawlers that collect material for training AI to stay off every page at portage.camp/@, and off the pictures and files people upload.

Search engines aren’t on that list, so a page whose owner has switched search on can still be found. Our own pages about Portage, and our summary at portage.camp/llms.txt, stay open to everyone.

The same pages and uploads also reserve text and data mining in machine-readable form, using the TDM Reservation Protocol, at portage.camp/.well-known/tdmrep.json.

It’s a request. A crawler that follows the rules stays out, and one that doesn’t can still read what anyone can see in a browser.

These are the crawlers our robots.txt names, and the line that keeps them off the pages:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /@

That’s all of it.

Questions about any of this? Email hello@cadenic.studio and a person will answer.

Make my page