Séimhe (sé / é)

  • 1 Post
  • 64 Comments
Joined 6 months ago
cake
Cake day: March 29th, 2026

help-circle

  • A lot of them use this in the HTML:

    <meta property="article:post_date" content="2026-09-24T15:42:36+0100">
    <meta property="article:post_modified" content="2026-09-24T15:54:02+0100">
    <meta property="article:published_time" content="2026-09-24T15:42:36Z">
    <meta property="article:modified_time" content="2026-09-24T15:54:02Z">
    

    I’ve scripted a fair bit of scraping in the past just by looking for those kinds of entries. But only focussed on a few dozen websites so I can’t say how common it is.

    If the fetching function can validate the data it should be able to continue without a date if needed.




  • I’m finding it hard to reconcile the people who find it useless with those who find it useful. My own experience with languages is that it speeds me up for tasks I have the skills to verify, but with ones I don’t it can slow me down doing the verification and fixes.

    I know software developers in a Fortune 500 company who swear that overall they are far more productive with “agents”. I don’t know if it’s a highly tailored system rather than the more all-purpose ones the general public have. So I do wonder if there are use cases for it.

    It’s ruining the planet and our societies, though. And it’s distracting us from what really matters. Net negative for sure.

    Edit: they’re also not very good at their most basic tasks. Researchers have found that when the major LLMs summarise articles, it misrepresents the facts 51% of the time.