theo-agent-dashboard

Unnamed repository; edit this file 'description' to name the repository.
Log | Files | Refs | README

theo_blog_scraped.json (61416B)


      1 [
      2   {
      3     "url": "https://theo.nya.pub/",
      4     "title": "Theo's Site",
      5     "date": null,
      6     "tags": [],
      7     "categories": [],
      8     "images": ["https://theo.nya.pub/wp-content/themes/personal-blog-theme/assets/images/hero.jpg"],
      9     "content": "This is my website. I am currently filling things out with various content of mine, and various projects I am working on.\n\n![Photo by Theo Jones](https://theo.nya.pub/wp-content/themes/personal-blog-theme/assets/images/hero.jpg)"
     10   },
     11   {
     12     "url": "https://theo.nya.pub/2023/02/02/ai-music-generation",
     13     "title": "Experimenting with AI Music Generation",
     14     "date": "2023-02-02",
     15     "tags": [],
     16     "categories": [],
     17     "images": [],
     18     "content": "I've been experimenting with AI music generation software lately and I have found it to be quite interesting. I've tried two programs, Mubert and Soundful, and I was pleasantly surprised with the results I got. Although the music generated wasn't very creative, it was good at replicating common background music styles. There was no noticeable distortion or excessively syllabic sounds in the audio, which I have seen in other services in the past.\n\nThe biggest issue with these AI music services is licensing. Most of the services I've tried claim copyright over the music generated using their tool and put many restrictions on how the music can be licensed back to the user. You can't use it in certain projects, or distribute the music on its own. Mubert has particularly restrictive licensing terms. To obtain full rights to the audio, the cost can be exorbitant, with Mubert charging upwards of $400, while Soundful, which has more reasonable terms, charges around $50. It almost seems like these services are trying to price their product just below what human artists would charge. This business model doesn't make much sense.\n\nWhen it comes to the ethics of AI art, I don't think it will harm artists as much as some people fear. For example, these AI music services I talked about can only replicate the most basic and generic types of music. They are still far inferior to human artists, even for creating background music for YouTube videos. I plan to get into streaming and making YouTube videos, and for that, I will still probably go with conventional music. I believe AI art will simply become another tool for creating art, reshuffling the deck a bit and potentially putting some people out of business or into business, but it won't be a sea change in the industry."
     19   },
     20   {
     21     "url": "https://theo.nya.pub/2022/06/01/berkeley-hp5-2022",
     22     "title": "Old Roll of HP5 — Berkeley, CA 2022",
     23     "date": "2022-06-01",
     24     "tags": [],
     25     "categories": [],
     26     "images": [
     27       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870030-1.jpg",
     28       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870033.jpg",
     29       "https://theo.nya.pub/wp-content/uploads/2026/05/000537870022-1.jpg",
     30       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870034-1.jpg"
     31     ],
     32     "content": "From an old roll of HP5 taken sometime in 2022 around Berkeley, CA and the San Francisco Bay Area, developed in 2023.\n\n![Berkeley Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870030-1.jpg)\n\n![Berkeley Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870033.jpg)\n\n![Berkeley Photo](https://theo.nya.pub/wp-content/uploads/2026/05/000537870022-1.jpg)\n\n![Berkeley Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000537870034-1.jpg)"
     33   },
     34   {
     35     "url": "https://theo.nya.pub/2022/12/10/best-camera-for-beginners",
     36     "title": "The Best Camera for Beginners",
     37     "date": "2022-12-10",
     38     "tags": [],
     39     "categories": [],
     40     "images": [],
     41     "content": "I believe that the best camera for beginners today is typically a fixed-lens bridge camera or a point-and-shoot. In the past, I have suggested to people that the first camera they purchase should be an entry-level DSLR or other interchangeable lens camera. However, it appears that the majority of camera makers are abandoning the sub-thousand-dollar interchangeable lens camera market in favor of the high end market.\n\nI recently used Canon's entry-level DSLR, the EOS T7, and it was not a positive experience. I experimented with it to get a feel for how it captured stills, how it took video, and everything else, but primarily I was curious as to whether or not it would make a suitable streaming camera to keep in a fixed spot and hook up as essentially a webcam utilizing Canon's webcam tool. And even compared to the experience of using some simple point-and-shoot cameras, it was a step down.\n\nIn many ways, the experience was inferior to that of a cell phone camera, and it felt antiquated. It also appeared as if the manufacturer developed the camera as a low-effort, entry-level product.\n\nIn addition, Canon is discontinuing its EOS M line of interchangeable-lens mirrorless cameras, which was formerly a very capable system. It is the first major camera platform I got into, but it appears like Canon won't really develop that many new cameras for that platform. I'm kind of bummed about that. They are moving on to the EOS R full frame system, which is more expensive, thus they have no plans to continue making lenses for that format.\n\nIn contrast, the market for budget-friendly point-and-shoot cameras has greatly improved with the introduction of optical image stabilization and computational photography features. A point-and-shoot camera with a one-inch sensor gives an excellent experience for a variety of everyday situations. It can perform well enough in low light that you can use it for the most of your daily tasks. Moreover, bridge cameras are becoming increasingly competent. Bridge cameras are, of course, fundamentally a more expensive market if you choose something with low-light performance. However, I believe that even the cheapest bridge cameras and super zoom cameras may produce decent results in the typical situations where one would use them.\n\nIn addition, there is a world of premium point-and-shoot or fixed-lens cameras that have also become quite good. So I sold my interchangeable lens system save for film cameras and switched to fixed lens cameras for most of my hobby work, because I believe it is a better deal these days."
     42   },
     43   {
     44     "url": "https://theo.nya.pub/2022/11/22/blogging-engine-research",
     45     "title": "Blogging Engine Research",
     46     "date": "2022-11-22",
     47     "tags": [],
     48     "categories": [],
     49     "images": [],
     50     "content": "I've been doing a bit of research to see what blogging engines exist that are in between WordPress (which is kind of a bloated mess) for running a small blog and Hugo and other static site generators which don't have web based UIs and a few other features.\n\nAn interesting one that's a minimalistic blogging engine but still not quite static site generator level minimalistic is Bludit. It looks like a very minimalistic blog that's just that doesn't have like a lot of extra features or bloat to it. It supports markdown.\n\nIt's not a static site generator but it has a flat file data structure to it so it's easy to backup because it doesn't have a MySQL database and there's not kind of that extra bloat of a MySQL database running on your server. That's either you have to tolerate more RAM usage or break containerization by using a kind of shared MySQL across all services on your server. So it looks like a pretty good option. I haven't replaced Hugo with it on my blog yet but it's a very interesting minimalistic blogging engine from what I can tell and how I've experimented with it a bit so far."
     51   },
     52   {
     53     "url": "https://theo.nya.pub/2026/05/10/caddy-in-smartos-native-zone",
     54     "title": "Caddy in a SmartOS Native Zone",
     55     "date": "2026-05-10",
     56     "tags": [],
     57     "categories": [],
     58     "images": [],
     59     "content": "On my home server, I am currently using Caddy as a reverse-proxy. For the public sites such as this Bookstack app, Caddy also provides SSL and other key security requirements.\n\nI have Caddy running in a native SmartOS zone. Caddy isn't that resource-intensive, so it could run on as little as 512 MB of RAM. In my zone, I gave it 2 GB of RAM. This is because I want to compile Caddy from scratch. I also gave it access to free CPU cores. I configured the zone in the SmartOS web UI instead of performing manual configuration.\n\nThe following are my steps to get Caddy running in a native SmartOS zone.\n\nFirst, install the Golang compiler.\n\n    pkgin update\n    pkgin install go122 git-base gcc12 gmake\n    \nYou might also want to install some creature comforts in your container, such as your preferred text editor. Personally, I like Nano, but you can pick what you want.\n\n    pkgin install nano\n    \nNext, set the environment variables to run the compiler.\n\n    export GOROOT=/opt/local/go122\n    export GOPATH=/root/go\n    export PATH=$GOROOT/bin:$GOPATH/bin:$PATH\n    \nTo make installing updates easier, you might want to make this persistent in your shell profile.\n\n    nano /root/.profile\n    source /root/.profile\n    \nNext, pull and install the Caddy source code.\n\n    go install github.com/caddyserver/xcaddy/cmd/xcaddy@latest\n    \nNext, build Caddy. You can add any extra plugins here using –with flags.\n\n    xcaddy build --output /opt/local/bin/caddy\n    \nCreate a user and group for Caddy.\n\n    groupadd -g 800 caddy\n    useradd -u 800 -g caddy -d /var/caddy -s /usr/bin/false caddy\n    \nCreate a folder for the Caddy server data and give the user and group permissions on it.\n\n    mkdir -p /opt/local/etc/caddy\n    mkdir -p /var/caddy/data\n    mkdir -p /var/log/caddy\n    \n    chown -R caddy:caddy /var/caddy /var/log/caddy\n    chmod 750 /var/caddy /var/log/caddy\n    \nCreate a Caddy file with the configuration options needed for your exact use cases.\n\n    nano /opt/local/etc/caddy/Caddyfile\n    \nCreate the SMF manifest needed to define the background service.\n\n    mkdir -p /var/svc/manifest/site\n    nano /var/svc/manifest/site/caddy.xml\n    \nImport and enable the service.\n\n    svccfg import /var/svc/manifest/site/caddy.xml\n    svcadm enable svc:/site/caddy:default\n    \nIf you change the Caddyfile, you can reload the Caddy configuration using the following command.\n\n    svcadm refresh svc:/site/caddy:default"
     60   },
     61   {
     62     "url": "https://theo.nya.pub/2022/02/01/censorship-degrades-public-trust",
     63     "title": "Censorship Degrades Public Trust",
     64     "date": "2022-02-01",
     65     "tags": [],
     66     "categories": [],
     67     "images": [],
     68     "content": "I recently watched the two controversial Joe Rogan episodes. Frankly, what I heard on those episodes was that out of the norm for current political discourse. I wasn't in anything close to hundred percent agreement with those guests had to say, of course.\n\nI would say the controversial guests did make good points. I would say the discussion was about 50% reasonable points (many of which haven't been discussed very much elsewhere) and about 50% crackpottery.\n\nI think censorship is counterproductive. When a large portion of the population is sympathetic to what your opponents say, you can't censor your way to public consensus and maintain public trust.\n\nAnd even if you could enforce the correct viewpoint on the public, censorship is fundamentally a Faustian bargain where society is creating extremely dangerous infrastructure of mass surveillance and social control. The type of centralized authority and centralized infrastructure that is required to force the censorship that we've seen recently on the Internet is intrinsically dangerous. It's an extreme act of hubris to think that if you give the right people that type of power, you'll get an utopia.\n\nAnother issue of widespread censorship is that suppressing discussion affects moderate speakers more than the extremists. Basically, when you're an extreme critic of policy you're going to get the ire of people no matter what you do. But when you're more moderate, people are going to tolerate you as long as you shut up on the parts where you disagree with the party line.\n\nAnd I think that there is a dynamic where people hear cognizant points from these speakers and when the discussion has been suppressed, they often first hear those cognizant points from the controversial people.\n\nThis gives a lot of credibility to the somewhat eccentric crackpots even when they are full of shit. I think that's why a lot of people are interested in hearing these controversial podcasts, podcasts like Joe Rogan's podcast are one of the rare places that you can hear actual discussion of some of these issues, instead of parroting of a party line is in many ways incoherent, arbitrary and rapidly changing.\n\nI don't think anyone thinks that it's the best possible source of commentary, I think a lot of people do think it's the only place where you won't hear commentary that's internally lockstep of everyone else.\n\nAnd the whole idea of building public trust by censorship and basically telling people that they're not allowed to have opinions on policies that are affecting their lives, is self-effacing fundamentally policymakers taking this approach will necessarily offend a large portion of the population and they will degrade public trust even more. It creates an adversarial relationship between policymakers and the public. And it creates a world where policymakers are too used to barking orders at the population, instead of finding ways to build public trust. This also ignores the vast variety of perspectives that are actually relevant to coming up with the best policy response. Different Americans are affected by the policies in different ways, policy elites will have their own biases and interests that are different than those of the typical American. This means that tight control over policy discussions will shut out many perspectives in a way that goes far beyond enforcing scientific objectivity or truth.\n\nI also think that as a whole US coronavirus policy has been highly corrupted by the fact that policymakers and media outlets who decided to treat China's response as the objectively ideal response, or at least a baseline for one response should look like, instead of the policy responses of Asian democracies like Taiwan or South Korea. The implication of this has been that US policymakers and media outlets have been acting like leaders from a communist dictatorship and using measures that would only reasonably be sustainable in an authoritarian state, instead of coming up with measures that would be reasonable for a pluralistic democracy."
     69   },
     70   {
     71     "url": "https://theo.nya.pub/2023/05/23/chatgpt-systems-admin-automation",
     72     "title": "ChatGPT Makes Automation Symmetrical with Doing",
     73     "date": "2023-05-23",
     74     "tags": [],
     75     "categories": [],
     76     "images": [],
     77     "content": "One of the clearest implications of ChatGPT for systems administrators is that it makes automating a task almost symmetrical with doing a task.\n\nOn the new file server I use for personal projects (dedicated server with a NVMe SSD boot drive and four hard drives as secondary file storage drives, I recently did a reinstall of Debian. I set this server up with the hard drives in a BTRFS raid 5. I installed Docker on it, and I set up the Apache web server to make the files on that server public. Cloudflare Tunnel was used to put that Apache server behind SSL.\n\nI took quick and rough notes on what commands were used and what may vary between servers and had ChatGPT create an automation script in Python.\n\nThe notes can be found here https://gist.github.com/theopjones/a7f2b6ba17f3de23826f688f0a87d01d\n\nThe prompt I used is\n\n> Create a python script to automate the server setup task in the following notes/log of a manual setup. Assume that the python script is running as root. In the case of commands which require manual intervention, wait for the user to conduct the manual intervention, the command should be started as part of the script.\n\nThe result of ChatGPT was the following, which is good enough to make this setup easily reproducible across servers or to document with code what the setup was so that it can be easily reproduced.\n\nhttps://gist.github.com/theopjones/6147770b550356e55d209e67549fb948"
     78   },
     79   {
     80     "url": "https://theo.nya.pub/2022/11/22/cheapest-gpu-for-ml",
     81     "title": "Finding Cheap GPUs for Machine Learning",
     82     "date": "2022-11-22",
     83     "tags": [],
     84     "categories": [],
     85     "images": [],
     86     "content": "I've done some investigation recently to try to figure out what's the cheapest GPUs around that would work for machine learning type tasks like running whisper or similar. I have a fairly beefy GPU in my computer, the A4000, which is an unusual configuration. It's a workstation GPU, not a consumer GPU. And it's a fairly high end GPU. I kind of got it because I mostly do productivity stuff on my computer, like photo, video, editing, some GPU intensive compute processes and things like that. But looking a bit into if there are lesser GPUs around just for recommendations to other people that would work. I think the obvious thing, and it's the one situation I tested with old equipment I have around would be the RTX 2060. It's kind of a consumer GPU, it goes used for about $200 from what I can tell on eBay and new for about $300 in a 12GB model. It's the cheapest consumer GPU that has high VRAM.\n\nAnd for most machine learning tasks that I'm interested in, VRAM is the limiting factor to an extent that's not true of gaming. On eBay I was able to find old workstation graphics cards that have a lot of RAM. One good example is the Nvidia M40, it has 12GB of RAM and I'm seeing it used for around $100. Like the absolute cheapest one that has enough RAM that I'm seeing is the K40, the Nvidia K40. And that also has 12GB of RAM. I would say the M40 would get pretty reasonable performance. The M40 has a pass mark score on GPU compute of 3775 operations per second. Kind of comparing that to the GPUs that I've kind of ran Whisper on, I guess it would do the large model in approximately one to one timing. One minute of audio input would take about a minute to process.\n\nThe GPU that I have, the A4000, gets about four to one. Four minutes of audio input would take a minute to process. The cheapest GPU that I've found that has enough VRAM has the $45 K40, has a pass mark score of around 2000 ops per second and that would like I think get like two minutes of processing time for like each minute of audio or maybe slightly worse than that. But I think like, I think there are a lot of kind of cheap GPU options if you're using the type of workflow that I use. And you just feed the speech tech software a pre-recorded recording and let it transcribe."
     87   },
     88   {
     89     "url": "https://theo.nya.pub/2023/10/15/disposable-camera-center-city",
     90     "title": "2023 Disposable Camera Roll — Center City Philadelphia",
     91     "date": "2023-10-15",
     92     "tags": [],
     93     "categories": [],
     94     "images": [
     95       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770002.jpg",
     96       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770014.jpg",
     97       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770001.jpg",
     98       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770023.jpg"
     99     ],
    100     "content": "Photos from various photowalks near Center City Philadelphia.\n\n![Center City Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770002.jpg)\n\n![Center City Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770014.jpg)\n\n![Center City Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770001.jpg)\n\n![Center City Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__1970__01__000538770023.jpg)"
    101   },
    102   {
    103     "url": "https://theo.nya.pub/2026/02/25/experience-using-opencode-latest-models",
    104     "title": "Experience Using Opencode on the Latest Models",
    105     "date": "2026-02-25",
    106     "tags": [],
    107     "categories": [],
    108     "images": [],
    109     "content": "I've been experimenting more with the latest LLM models for coding. And it's pretty impressive how far things have come, and how these tools are pretty impressive.\n\nI've mostly been using the Kimi 2.5 model with Opencode as the coding agent. I still find that mix pretty great. I think the whole vibe coding/AI assisted programming workflow that Opencode and similar encourage might not be the best for quality code, but it is pretty addictive seeing that type of rapid progress. Until you get to the very highest (expensive) tier of Claude/Anthropic and OpenAI models, Kimi performs basically on par or better than what the biggest companies offer.\n\nAnd these coding agents can take care of a lot of the boring drudgework of programming. They are good enough right now that I don't have to spend too much time manually intervening and fixing what the LLM did — these tools are getting pretty accurate.\n\nI've spent many hours working on a project, tweaking it back and forth, with the main thing stopping me from spending even more time is the fact that I have to pay for the credits to run inference for the models. Once you spend the time working through all of quirks these tools have gotten pretty smooth as far as workflow goes. It's genuinely _fun_ to do this on the recent LLM models that have come out.\n\nCost right now is the only real problem, these things will burn through tokens by the millions. It's pretty clear that the $20/mo for coding agents tier from OpenAI etc (even though the limits are being tightened) are being subsidized pretty aggressively. When you compare the amount of time you get on a coding agent from such a plan vs what open source alternatives cost, OpenAI etc can't be making money from the coding agent offerings. On the other hand, it's probably cheaper to use a self hosted frontend (OpenWebUI etc with a hosted inference API) than it is to pay for a paid tier of ChatGPT.\n\nI've also noticed that Opencode and other open source agents/frontends are very sensitive to token output performance. Using a somewhat more expensive inference provider that provides fast performance will improve the experience quite a bit. Switching API providers basically fixed some of the issues I was having with the model freezing etc.\n\nThe project I've been working on as part of my testing is this https://git.selfhosted.onl/theo/marginleaf\n\nIt's a personal blogging CMS. It can do the typical blogging engine things, but instead of a frontend editing interface, I created an API, and I built some tools that allow me to fully manage it from Open WebUI, which opens some pretty neat possibilities. It feels like a somewhat interesting possibility to have some of these chat tools getting good enough that they can be the main interface with an application — instead of a more traditional web UI.\n\nIn particular, the Open WebUI tools can be found here https://git.selfhosted.onl/theo/marginleaf/src/branch/main/openwebui_tools\n\nBut mostly I created it because it's fun to work on that type of thing."
    110   },
    111   {
    112     "url": "https://theo.nya.pub/2022/12/11/experimenting-with-owncast",
    113     "title": "Experimenting with Owncast",
    114     "date": "2022-12-11",
    115     "tags": [],
    116     "categories": [],
    117     "images": [],
    118     "content": "Experimenting with Owncast – an open source twitch-like streaming application.\n\nFirst stream will probably be sometime this evening.\n\nBecause of the potentially high bandwidth usage, I'm setting it up on a VPS that has higher internet bandwidth than anything I use. This VPS will probably get used for other livestreaming/messaging tasks, and various other things I don't want to be 100% dependent on my home internet connection."
    119   },
    120   {
    121     "url": "https://theo.nya.pub/2022/12/07/gpt3-creative-script",
    122     "title": "Using GPT-3 for Creative Writing",
    123     "date": "2022-12-07",
    124     "tags": [],
    125     "categories": [],
    126     "images": [],
    127     "content": "Created another script that uses a single run of the edit mode of GPT-3 with a high temperature (ie. giving GPT-3 a high degree of creativity). But it runs it three distinct times.\n\nhttps://gogs.theopjones.blog/theo/LittleScripts/src/master/transcribefoldermultiple.py\n\nThe results are interesting.\n\nUnedited Transcript:\n> I've been experimenting a bit with using GPT-3 to process speech-to-text transcripts, which in their raw form contain no line breaks, no paragraph breaks, kind of off-text because it's a direct transcription of my speech, like not how I would normally write it. I'm feeding, and I have a little Python script written to feed these raw, unprocessed speech-to-text transcripts into GPT-3. Of course, GPT-3 can't be ran locally, so it has to make an external API call. But how the script I wrote works is it makes one API call to have the text split up into individual paragraphs, and it makes another set of API calls for each paragraph to correct the grammar, style, spelling, and all of that. I did the two-part thing because based on my experimentation, GPT-3 doesn't really like being given a huge wall of text, so splitting it up into paragraphs is one of the best techniques I found to get GPT-3 not to remove a lot of text without creating replacement text or add totally new text. From what I can tell, the little script I wrote is able to keep things pretty faithful to how I originally dictated while still punching up the grammar and resolving a lot of the editing I would have to do to make a speech-to-text transcript usable on my blog or something. So I think it's helpful because it reduces a lot of really error-prone stuff that comes with using speech-to-text to write. I've uploaded a little Python script. I've used use slash created, and you can find it below.\n\n(Three GPT-3 runs with increasing creativity were shown as examples of the output.)"
    128   },
    129   {
    130     "url": "https://theo.nya.pub/2022/12/06/gpt3-speech-to-text",
    131     "title": "Processing Speech-to-Text with GPT-3",
    132     "date": "2022-12-06",
    133     "tags": [],
    134     "categories": [],
    135     "images": [],
    136     "content": "I've been experimenting with using GPT-3 to process speech-to-text transcripts. These transcripts, in their raw form, contain no line breaks or paragraph breaks, and are not how I would normally write because they are direct transcriptions of my speech. I have written a small Python script to feed these unprocessed transcripts into GPT-3. Of course, GPT-3 cannot be run locally and requires an external API call.\n\nBut how the script I wrote works is that it first makes one API call to split the text into individual paragraphs, and then it makes another set of API calls for each paragraph to correct the grammar, style, and spelling. I opted for the two-part approach because, based on my experimentation, GPT-3 doesn't really handle large blocks of text very well. So, splitting it up into paragraphs is one of the best techniques I've found to prevent GPT-3 from removing too much text without creating replacement text or adding totally new text.\n\nFrom what I can tell, the small script I wrote is able to keep things faithful to how I originally dictated, whilst still improving the grammar and resolving much of the editing I would have to do to make a speech-to-text transcript usable on my blog or something. Thus, I think it's helpful as it reduces a lot of the error-prone aspects associated with using speech-to-text to write.\n\nThe script can be found here https://gogs.theopjones.blog/theo/LittleScripts/src/master/transcribefolder.py\n(this post is just the output of this workflow, with minimal additional editing)"
    137   },
    138   {
    139     "url": "https://theo.nya.pub/2022/09/17/home-servers-tunneling-etc",
    140     "title": "Home Servers, Tunneling, etc",
    141     "date": "2022-09-17",
    142     "tags": [],
    143     "categories": [],
    144     "images": [],
    145     "content": "As a follow-up to my post earlier this week, I'll discuss some other interesting things about setting up a home server.\n\nUnfortunately, the technology here is a little bit opaque, and I'm not really aware of any good documentation that exists on how to set up servers that is newbie friendly. Most of the writing here doesn't really start from first principles, and a lot of what you'll find is aimed at super knowledgeable people, or people like IT systems administrators.\n\nThere's a lot of stuff on Internet forums and on Reddit and on various peoples blogs. And when I figure this stuff out I do a lot of Googling and visiting Reddit threads, and visiting Stack Overflow threads.\n\nI've thought about writing a bit more about how the technology works and how to set up this type of server. But this is not something i've done yet\n\nSetting up HTTPS has become a lot easier than it used to be. Caddy, which I use for the reverse proxy in my server basically handles SSL without me having to do much. There's also a helper for NGINX which deals with a lot of the setting up the reverse proxy and setting up SSL.\n\nThe existence of Let's Encrypt has basically eliminated the need to buy SSL certificates from designated certificate authorities, and it's what the tools I mentioned above are built on top of.\n\nThe security situation is kind of a mixed bag, there are some tools I ran into that have super insecure default configurations, fortunately the security of the most common software programs has improved a lot compared to where it used to be. Most of the big tools that you'll run into like Web servers and so on are pretty much secure by default, you would have to actively change the configuration in undesirable ways to make it insecure.\n\nAnd I think container programs like Docker and so on also help a lot with security, basically every application I have running on my server has its own docker container. The Caddy reverse proxy works as the glue between these containers.\n\nDocker is a way of packaging software programs with needed libraries and dependencies, it functions in a very VM like way – there is a high level of isolation between the different containers by default. This isolates security issues, if one of the services running on the server gets owned it's hard for the hacker to privilege escalate to the rest of the server, so it's possible to just deal with the security issue by just nuking that one container and starting fresh.\n\nAdditionally, since there are a lot of docker images that are packaged either by the developers of the software or by someone else upstream, it's pretty easy to find a docker container where everything's packaged into a pretty secure by default container.\n\nFor backups, I use the Duplicati tool, set to make daily backups of the server. It's possible to back up to a portable hard drive, or to another server with Duplicati on it that's off-site. I haven't taken any of these purist paths, and I have taken the more non-self hosted route of uploading my data to a cloud storage provider (in this case Wasabi).\n\nDuplicati is capable of encrypting the backups before they go to the cloud storage provider, or friend's server, or whatever else you're using for your remote backup.\n\nThere are two ways to connect the server to the outside world.\n\nThe traditional way, what I used, is to get a static IP address from your ISP. AT&T, who I use for my Internet, sells static IP addresses in a /29 block, that is six usable IP addresses, unfortunately, they won't give you just one static IP address. Additionally, I still have access to one dynamic IP address from them.\n\nMy router/gateway/modem gets assigned one of the static IP addresses, the home server gets assigned another, basically every other device on my network gets put behind the dynamic IP address.\n\nThe more newfangled way of connecting your server to the internet is to use a tunneling service.\n\nNgrok and PageKite are two pretty good examples of these types of services. Your server opens a connection to the tunneling service, and the tunneling service assigns an IP address to your traffic (or subdomain that can be attached to a domain name as a CNAME record).\n\nThe one I've done the most experimentation with has been Cloudflare Tunnel. The biggest problem with this service is that it kind of adds another ISP-like intermediary between your server and the user. This is a step back in terms of avoiding over dependence on centralized services, but since the data itself lives on a server you control, it's still an improvement over your standard content silos or proprietary services.\n\nCloudflare goes a bit further than many of the other tunneling services in terms of the amount of integration with your site – and not only routes the data, but also takes over the SSL certificate and does a lot of filtering and analysis on the traffic.\n\nCloudflare tunnel is probably the option I'd recommend to people who don't have a super in depth technical knowledge.\n\nIt handles reverse proxying, it can put private services behind an authentication portal, it provides DDOS protection, a content delivery network, rate limiting for bots, and a web application firewall."
    146   },
    147   {
    148     "url": "https://theo.nya.pub/2023/05/23/im-looking-for-work",
    149     "title": "I'm looking for work",
    150     "date": "2023-05-23",
    151     "tags": [],
    152     "categories": [],
    153     "images": [],
    154     "content": "I was recently laid off from my previous company.\n\nI'm a seasoned IT and customer service professional with over five years of experience. My skills extend from software deployment and support to Linux administration and Python scripting for automation.\n\nI've acted as an administrator for major SaaS platforms such as Google Workspace, Docusign, email marketing tools (PersistIQ, ActivePipe), CRMs (CopperCRM, Contactually, Follow Up Boss), and Okta, effectively resolving email infrastructure issues. Also, I've offered on-call and after-hours support for urgent user requests.\n\nMy proficiency in open-source platforms includes managing LAMP + Nginx servers, working with cloud compute/VPS hosting platforms, and utilizing Linux for desktop and server projects. I have automated tasks using Python and other scripting languages, focusing on account creation, data migrations, and infrastructure management. Additionally, I've used low-code platforms like Zapier, and have some familiarity with the Dell Boomi Platform.\n\nOne noteworthy accomplishment is automating most of the user onboarding process, allowing accelerated growth without increasing IT staff. I've also efficiently transitioned data from one CRM system to another, leveraging APIs to rebuild account environments.\n\nI am adept at defining requirements with software engineers and vendors for new product rollouts. I am well-versed in IT security, including implementing and documenting new security processes and mitigating threats.\n\nMy experience with support ticketing and project management systems spans Service Cloud, atSpoke, Jira, and Asana.\n\nFurthermore, I hold degrees in Geography and Ecology and Evolutionary Biology from the University of Arizona, with a focus on geographic information systems. I've tutored STEM and geography subjects and have experience in GIS and scientific data analysis from internships.\n\nMy desired salary for a new role is $85,000/yr, though I'm open to $60,000-$85,000 depending on the total compensation package, the nature of the employer, and the status of my other interviews. While I prefer a W2 role, I'm also open to contract-to-hire and independent contractor status, and am available for freelance work that doesn't conflict with full-time employment.\n\nFor more information, please reach out to me by email tjones2@fastmail.com or through my LinkedIn profile https://www.linkedin.com/in/theodore-jones-7b89b7269/"
    155   },
    156   {
    157     "url": "https://theo.nya.pub/2026/02/20/kimi",
    158     "title": "Kimi 2.5 and Self-Hosting Open WebUI",
    159     "date": "2026-02-20",
    160     "tags": [],
    161     "categories": [],
    162     "images": [],
    163     "content": "Been poking around with the Kimi 2.5 LLM and also started self-hosting Open WebUI on my server (a self-hosted ChatGPT-style web frontend for LLM APIs).\n\nKimi probably isn't the best model on the market, but Kimi 2.5 is the first time I've used a truly open source model that feels to be vaguely in the same category of performance as ChatGPT, etc. And I don't really feel much of a penalty using it vs ChatGPT.\n\nOf course, running it directly is _way_ beyond what any device I have can do reasonably well.\n\nBut there are already API providers around offering it with very favorable privacy and data retention policies, so I'm probably going to switch to using it over ChatGPT.\n\nI wouldn't recommend using the chat/API offered by the model's creator–I don't really trust that company.\n\nIf I self-host the front end, all of the actually sensitive data like chat logs etc are stored on my server.\n\nOpen WebUI is pretty cool. It works almost as well as ChatGPT does. I've run into some issues with the model occasionally freezing during processing, but I've occasionally seen that type of thing with other LLM providers.\n\nIt has a search integration that works with the model so it can web search etc. It's pretty customizable.\n\nI quickly created a custom tool that the model can use which queries the OpenAlex API to find open access academic articles. The code for that can be found here https://git.selfhosted.onl/theo/openwebui-tools-skills/src/branch/main"
    164   },
    165   {
    166     "url": "https://theo.nya.pub/2023/10/20/philadelphia-chinatown",
    167     "title": "Philadelphia Chinatown (2023 Oct)",
    168     "date": "2023-10-20",
    169     "tags": [],
    170     "categories": [],
    171     "images": [
    172       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010864.jpg",
    173       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010870.jpg",
    174       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010876.jpg",
    175       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010859.jpg",
    176       "https://theo.nya.pub/wp-content/uploads/2026/05/47887819-scaled.jpg",
    177       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010881.jpg",
    178       "https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010883.jpg"
    179     ],
    180     "content": "While taking these photos, I saw a lot of signage about a proposed stadium for the 76ers.\n\nMost of this was in opposition (I am not informed enough to give a direct opinion regarding the issue).\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010864.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010870.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010876.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010859.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/47887819-scaled.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010881.jpg)\n\n![Chinatown Photo](https://theo.nya.pub/wp-content/uploads/2026/05/PhotoLibrary__2023__10__L1010883.jpg)"
    181   },
    182   {
    183     "url": "https://theo.nya.pub/2026/01/20/pinning-footer-to-bottom",
    184     "title": "Pinning Footer to Bottom of Page in Bootstrap Studio",
    185     "date": "2026-01-20",
    186     "tags": [],
    187     "categories": [],
    188     "images": [],
    189     "content": "Some of the themes that come with bootstrap studio don't have the footer pinned to the bottom of the page.\n\nThe instructions in this link are helpful here: https://forum.bootstrapstudio.io/t/footer-always-at-the-bottom-of-the-page/7517\n\nBasically, create custom sitewide CSS (ie. through a .css file under the styles folder of the design), with the following:\n\n    body {\n        display: flex;\n        flex-direction: column;\n        height: 100vh;\n    }\n    footer {\n        margin-top: auto;\n    }"
    190   },
    191   {
    192     "url": "https://theo.nya.pub/2022/04/21/quadratic-voting-does-not-scale",
    193     "title": "Quadratic Voting Does Not Scale",
    194     "date": "2022-04-21",
    195     "tags": [],
    196     "categories": [],
    197     "images": [],
    198     "content": "Quadratic voting is a potential voting method that has gotten a fair amount of discussion in various places, one of the most notable presentations on this is in Radical Markets. While the game theoretic justification for this voting method is sound under optimal conditions, with low information/transactional costs, and perfectly rational actors, I believe that there are flaws in this idea that make it unusable in most real world circumstances where it is being proposed. It is a system that is perfect on paper, but unsuited to the real world.\n\nA flaw of many real world voting systems is that there is not a good way to allow voters to provide information about the relative importance of issues. This means that people who only have a weak preference on an issue will be in effect over represented in political outcomes on that issue. Quadratic voting is a proposal to fix this issue.\n\nIn a QV ballot a voter has a number of points that they can allocate across issues. Allocating more points to an issue makes the vote on that issue weighted more. The value of each point declines as you add more points to an issue. Accordingly, there is an incentive to split your points across multiple issues.\n\nWhat do I see as the problems with this proposal?\n\nIn summary, it runs into issues with very large elections, and breaks down when people don't act as Homo economicus purely rational self-interested actors.\n\nTwo examples: How to optimally allocate voting points — The optimal strategy in a QV system would be to allocate votes in proportion to the value of a vote, not the subjective importance of the outcome of that particular issue. This can result in counterintuitive allocations when both large and small elections are on the same ballot.\n\nHigh minimum threshold for issue importance — Quadratic funding/voting can only really be used for very huge issues and in fairly small communities given the practical limits of how people think, defeating one of its main features."
    199   },
    200   {
    201     "url": "https://theo.nya.pub/2022/11/23/response-to-mastodon-discussion",
    202     "title": "Thoughts on Mastodon",
    203     "date": "2022-11-23",
    204     "tags": [],
    205     "categories": [],
    206     "images": [],
    207     "content": "In response to a Tumblr post asking about Mastodon.\n\nThe short and quick answer is that Mastodon is an open source program that provides Twitter like functionality. It's something that you can use to set up a social media website of your own.\n\nIt is possible for different instances of Mastodon to talk to each other but this highly depends on how the particular administrators have their instances configured and it is fairly common for instances to refuse to communicate with each other often for trivial reasons or just the administrator's personal preference.\n\nSo I would call Mastodon at best a semi-decentralized system because the general assumption of Mastodon is that most users will join an instance that's ran by someone else and that most users won't run their own instance. There is very limited portability of accounts between instances. Identity on Mastodon is completely tied to the individual instance.\n\nIt is possible to run your own instance just for you, and get some other instances to talk to your instance, but most people use other instances. The software isn't really built for one user instances, and generally assumes that an instance has a lot of users. Managing a Mastodon instance is relatively complicated compared to a lot of other server software.\n\nMastodon instances usually have much heavier handed moderation than other social media.\n\nMy take on Mastodon is fairly negative. I think it's a system that somehow manages to reproduce kind of worst of Twitter but also has none of the benefits of true decentralization.\n\nMost of what Mastodon is good at can be fundamentally done by other ways. My opinion on this type of thing is that the protocols of the old open blogosphere fundamentally worked. The reasons why the old blogosphere kind of died out are unrelated to the things that Mastodon is optimizing for."
    208   },
    209   {
    210     "url": "https://theo.nya.pub/2022/09/12/self-hosting-text-website",
    211     "title": "Self-Hosting a Text Heavy Website is a Solved Problem (Even on a Home Server)",
    212     "date": "2022-09-12",
    213     "tags": [],
    214     "categories": [],
    215     "images": [],
    216     "content": "A while ago I got a mini PC and turned it into a home Web server. This turns out to be a remarkably effective way to host a website. The blog you're reading right now is hosted on on this mini PC. And this home server is connected to my standard home Internet. And it's not just my website, I host a lot of other services in order to improve the privacy of my data.\n\nI decided to do some testing to figure out how much traffic my set up can handle.\n\nThe Server: A fairly inexpensive Beelink mini PC with 8 GB of RAM and a 256 GB mSata SSD, roughly $170.\n\nThe Software: Debian Linux with the Caddy web server, most services in Docker containers.\n\nThe Internet Connection: 1 Gbps symmetrical fiber, with static IP addresses.\n\nThe Website: Built with Hugo static site generator.\n\nResults: My home server can handle about 300 HTTPS requests per second. The limiting factor is the server's ability to handle simultaneous requests. Internet bandwidth didn't seem to matter much (didn't exceed 50 Mbps during testing).\n\nThis is enough to withstand getting posted on the front page of Reddit (the 99th percentile load from Reddit front page is about 360 requests per second).\n\nMy conclusion is that a self hosted blog that is well optimized can be hosted on a standard home Internet connection using a cheap computer as a server. Hosting text heavy content in a decentralized way is basically a solved problem."
    217   },
    218   {
    219     "url": "https://theo.nya.pub/2023/03/27/setting-up-goblog-on-freebsd",
    220     "title": "Setting up GoBlog on FreeBSD",
    221     "date": "2023-03-27",
    222     "tags": [],
    223     "categories": [],
    224     "images": [],
    225     "content": "GoBlog is a blogging engine that I have used on my personal blog, and various other personal projects. I'm going to do a walkthrough of how to set this up on a FreeBSD server.\n\nIf you want a quick TLDR, here is a shell script that automatically spins up GoBlog: https://gist.github.com/theopjones/e09c9713c10f4000d154de50c438d2ba\n\nIt's a blogging engine with fairly few users, but for the technically inclined, it makes a good personal blog. It is very performant and supports a lot of interesting social features, including most of the IndieWeb standards. It can also talk to Mastodon and other ActivityPub services.\n\nSteps include: installing go-devel git gcc sqlite3 bash packages, cloning the GoBlog source from Git, compiling, creating config files, and setting up an RC script for FreeBSD service management.\n\nFull installation script: https://gist.github.com/theopjones/e09c9713c10f4000d154de50c438d2ba\nConfig generator script: https://gist.github.com/theopjones/748c296b3c33881352bb7ac72772ae67\nRC script: https://gist.github.com/theopjones/d62e480a71f5cbcead7e381ffd422fda"
    226   },
    227   {
    228     "url": "https://theo.nya.pub/2023/02/02/some-thoughts-on-the-ethics-of-ai-artgenerative-ai",
    229     "title": "Some thoughts on the ethics of AI art/generative AI",
    230     "date": "2023-02-02",
    231     "tags": [],
    232     "categories": [],
    233     "images": [],
    234     "content": "AI art is getting a lot of controversy for its implications for current artists. I for one think that some of the fears of human artists getting fully displaced by automation is a bit over stated. I think it will be just another tool that's used to create art.\n\nHowever, what worries me the most about the increasing role of AI tools is their closed nature. Currently, the AI models and their outputs and inputs are owned by just a few companies, leaving most users locked out.\n\nI have a strong concern that this will concentrate the art market, displacing the decentralized infrastructure and ecosystem of small business artists with a much more centralized art world.\n\nThe majority of the significant recent generative AI models are proprietary. Even source-available models like Stable Diffusion are not fully free and open source due to \"toxic candy models\" concerns.\n\nA concerning trend is the RAIL (Responsible Artificial Intelligence Source Code) License, which imposes restrictions on how users can use the output generated by the tool. This is a departure from the open-source community's consensus.\n\nAI art is just another method of art that uses technology to probe and sample an extrinsic space outside of the artist's mind, similar to how photography creates art by sampling from the physical environment.\n\nUsing copyright as a means to control AI output is dangerous. If this mindset spreads, it would rewrite the balance of power between software companies and consumers. Software is a functional work — control over software used to make art is fundamentally exerting control over a method or technique. The freedom to use and modify software is critical."
    235   },
    236   {
    237     "url": "https://theo.nya.pub/2025/01/20/stenomasks-and-speech-to-text",
    238     "title": "Stenomasks and Speech to Text",
    239     "date": "2025-01-20",
    240     "tags": [],
    241     "categories": [],
    242     "images": [],
    243     "content": "For a while I've had this StenoMask thing, which is a sound isolated box that can be talked into for speech recognition. I think the notational thing it's commonly used for is court reporters speaking into it for notes that can be transcribed later. Of course, my use case with it is writing without a keyboard and similar.\n\nWhen I first started experimenting with it, I found that it was really hard to get any kind of acceptable accuracy with speech recognition software.\n\nI've been trying it again now. Speech recognition software has gotten to the point where I can talk to it normally and it basically just works when transcribing. Which makes the thing actually useful for me now.\n\nThis is what I am using: https://whispertyping.com/\n\nIt would be interesting to try to give Dragon NaturallySpeaking a try again. It's what psychologists have recommended for me for some of the relevant disabilities I have. Dragon is very expensive (hundreds of dollars), so doesn't feel worth it to give it another try."
    244   },
    245   {
    246     "url": "https://theo.nya.pub/2023/03/23/this-blog-is-back-up",
    247     "title": "This blog is back up",
    248     "date": "2023-03-23",
    249     "tags": [],
    250     "categories": [],
    251     "images": [],
    252     "content": "This blog is back up.\n\nIt was down for a while after I moved apartments.\n\nHad to move it to an external server, because 1) the computer I was using as a home server got shipping damaged during the move, and 2) my new home internet has much slower upload speeds than the fiber connection in my last apartment."
    253   },
    254   {
    255     "url": "https://theo.nya.pub/2022/01/30/thoughts-on-cryptocurrency-and-web-30",
    256     "title": "Thoughts on Cryptocurrency and Web 3.0",
    257     "date": "2022-01-30",
    258     "tags": [],
    259     "categories": [],
    260     "images": [],
    261     "content": "I'm going to provide some thoughts on cryptocurrencies, NFTs, and the concept of a \"web 3.0\". I'm not particularly enthusiastic about a lot of the things under that umbrella, I think the technology as exists today either has fundamental flaws in many cases or is a solution in search of a problem.\n\nFundamentally, these technologies are basically ways to replace centralized intermediaries. The fundamental technical problem is that these decentralized protocols will inherently be extremely inefficient compared to centralized alternatives.\n\nThe use case of these technologies will be naturally limited to cases where it's worse to trust an intermediary — things like transactions that are discouraged by mainstream banks. The system of mass surveillance that the modern banking system creates is really a net negative for society.\n\nThere are fairly plausible systems for digital cash that ultimately involve normal intermediaries like banks. This technology feels fundamentally better.\n\nIPFS is vastly slower than HTTP. The overhead for searching for files is huge. My experimentation came to the conclusion that the performance is so abysmal that its nowhere near possible to host a website on IPFS reliably.\n\nArt NFTs are basically useless. It's pure artificial scarcity and speculation.\n\nDespite the fact I think the technology feels like fundamentally a dead end, I feel somewhat sympathetic when I see a lot of critics coming from the perspective of authoritarianism."
    262   },
    263   {
    264     "url": "https://theo.nya.pub/2022/03/27/thoughts-on-mastodon-and-avoiding-content-silos",
    265     "title": "Thoughts on Mastodon and Avoiding Content Silos",
    266     "date": "2022-03-27",
    267     "tags": [],
    268     "categories": [],
    269     "images": [],
    270     "content": "I recently set up a Mastodon instance (username @theo@theopjones.com)\n\nThe 500 character text limit on Mastodon does seem a lot better than Twitter's shorter character limit.\n\nIn my ideal world, most people on the internet would use open source software running on commodity infrastructure. I want a world in which where you decide to host your content fundamentally doesn't matter and there is competition in content hosting.\n\nThe big thing that worries me about Mastodon from a structural perspective is that it simultaneously is generally structured in a way that means the vast majority of users won't run their own instance AND has a primary mode of content moderation built around instance administrators blocking other instances. This could easily replicate the situation with email where there are very much first-tier email hosts.\n\nIf I were designing things, content filtering would be user based — split into three categories: people the user has directly opted into seeing, people trusted by someone the user trusts, and the rest of the world. The rest of the world would have to do something costly like a digital stamp or proof of work.\n\nI like what I see in the Indieweb project. The emphasis on personal domains instead of shared instances is good. But all of the actually existing software here has been a pile of half-working kludges in my testing. While Mastodon actually works."
    271   },
    272   {
    273     "url": "https://theo.nya.pub/2022/11/22/tumblr-feed-thoughts",
    274     "title": "Switching Away from Apple",
    275     "date": "2022-11-22",
    276     "tags": [],
    277     "categories": [],
    278     "images": [],
    279     "content": "I am talking while going through my feed on Tumblr.\n\nThe first interesting post is about switching away from macOS. I recently actually switched away from macOS myself. A lot of the reasons why I have been switching away from Apple products is that Apple's business practices have become a lot worse. I bought a proper workstation desktop and put Linux on it. I sold my Macbook and switched to Linux and Windows.\n\nThe biggest issues I ran into with the transition are device incompatibilities — specialized devices like a sound recording DAC paired to Thunderbolt ports or macOS software. Also dependency on proprietary software, particularly Photoshop and Lightroom.\n\nFortunately Adobe is starting to have a really good web app so it's possible to just use Lightroom as I normally did.\n\nI recently bought a Google Pixel Android phone and moved my phone plan over. That will be the last big Apple device.\n\nWhat Apple's doing that's kind of new in its badness is the extent to which Apple doesn't let you treat your device as your device and tries to block what you can do with it. When Apple pressured Tumblr into blocking certain content because their App Store refused to accept the Tumblr app. That's novel. The fact that Apple uses the fact that your device is locked to them and you can't sideload apps. It's a threat to software freedom that's new."
    280   },
    281   {
    282     "url": "https://theo.nya.pub/2022/02/05/urban-development-costs",
    283     "title": "The Costs and Benefits of Urban Development are Distributed Very Unequally",
    284     "date": "2022-02-05",
    285     "tags": [],
    286     "categories": [],
    287     "images": [],
    288     "content": "Urban development in big cities is very controversial, and there are politically powerful movements that oppose almost all new construction in big cities.\n\nThere is one explanation that in my opinion just doesn't hold water — that incumbent property owners want to increase property values. The reason this doesn't hold water is that urban cores with the most opposition to development (San Francisco, NYC) have some of the lowest rates of property ownership.\n\nMy explanation is based around the fact that the costs and benefits of development are distributed very unequally. Incumbent renters have very few personal gains from new development except under very long time frames.\n\nA 2019 article discusses the economic implications of development: increased density brings higher wages, more job opportunities, more innovation, less car dependency, and better public transit. However, property values will be increased, which is great if you own property but not if you rent.\n\nA 2015 paper concluded that wages would increase drastically on a national scale, and the economy would've grown 50% more between 1964 and 2009 if zoning regulations were more permissive.\n\nThe distributional impacts of new housing construction cannot be ignored and are the primary source of opposition to new development. Housing development must be done in a way that makes sure that the typical American benefits."
    289   },
    290   {
    291     "url": "https://theo.nya.pub/2022/09/24/voice-typing-wrapper-around-whisper",
    292     "title": "Voice Typing Wrapper Around Whisper",
    293     "date": "2022-09-24",
    294     "tags": [],
    295     "categories": [],
    296     "images": [],
    297     "content": "I just wrote a voice typing wrapper around Whisper. It types what I say as keyboard input, and it creates a system tray icon to turn on and turn off the dictation.\n\nhttps://github.com/theopjones/voice-typing\n\n(I just created it, it might have bugs, only tested on Linux)\n\nI'm not sure how much additional time I want to invest in this little project. Because I'm not an expert in this type of technology or AI in general.\n\nI think right now I have something that's a very interesting proof of concept. But while testing it, I have encountered a few bugs and little glitches. And I definitely don't get the same exact level of accuracy while voice typing with this tool that I'd get just pre-recording my voice and feeding it in all at once.\n\nInternally what this does is it breaks up the audio into small little snippets and parses each one of those snippets automatically. This doesn't do wonders for interacting with the underlying model because it's not consistent with the assumptions being made in Whisper.\n\nThe underlying Whisper model assumes that it is dealing with really long blocks of audio. When I dictate a lot to it at once it kind of jumbles up the grammar/punctuation.\n\nIn its current state, it comes pretty close to meeting my immediate need for a dictation program/voice typing program."
    298   },
    299   {
    300     "url": "https://theo.nya.pub/2022/09/23/whisper-speech-to-text",
    301     "title": "The Whisper Speech to Text Library Appears Really Powerful",
    302     "date": "2022-09-23",
    303     "tags": [],
    304     "categories": [],
    305     "images": [],
    306     "content": "There's a new speech-to-text program/library that just got released by OpenAI as open source called Whisper and it's impressed me quite a bit so far. It's really powerful and it competes pretty well with the incumbent major speech-to-text tools in terms of accuracy.\n\nThe caveat being that its not a full featured tool. Currently all it does is convert an audio file to text. It's a command line tool so far. It doesn't have anything more sophisticated like simulated keyboard input or training.\n\nThe accuracy is better than even mature speaker dependent systems like Dragon. It has a very strong model of grammar and gets things that are really difficult for most speech to text programs like capitalization, prepositions, or small words. It gets a lot of technical/specialized terms right.\n\nIt has the same accuracy to expect from a speaker dependent program that's been trained a while on your voice even though it's a speaker independent program.\n\nI also tried a stenomask (a microphone that goes right up against your mouth for privacy). Accuracy declined but was still pretty impressive and quite usable.\n\nIt's only kind of sort of open source — you can download the tool and a pre-built model, but software to generate that model from audio hasn't been released yet. The model is based on not open source licensed data.\n\nI have been long interested in speech-to-text systems because I have a handwriting disability that makes it hard for me to quickly type and write normally."
    307   }
    308 ]