Misunderstanding Computers

Why do we insist on seeing the computer as a magic box for controlling other people?
人はどうしてコンピュータを、人を制する魔法の箱として考えたいのですか?
Why do we want so much to control others when we won't control ourselves?
どうしてそれほど、自分を制しないのに、人をコントロールしたいのですか?

Computer memory is just fancy paper, CPUs are just fancy pens with fancy erasers, and the network is just a fancy backyard fence.
コンピュータの記憶というものはただ改良した紙ですし、CPU 何て特長ある筆に特殊の消しゴムがついたものにすぎないし、ネットワークそのものは裏庭の塀が少し拡大されたものぐらいです。

(original post/元の投稿 -- defining computers site/コンピュータを定義しようのサイト)

Sunday, August 15, 2021

The Real Reason IBM Chose the "Wrong" CPU for the PC

[JMR202108181958 -- adding summary:

My own sisters don't want to read this because I'm speaking too much Geek. Re-reading it, I guess they're right. And what I wrote seems to wander all over without apparent reason -- not apparent unless you already know what I'm trying to say.

So, I guess I should put a high-level summary up front here:

Up until twenty years ago, common wisdom was that IBM picked the "right" CPU for the IBM PC. Only crackpots like me thought differently. But the evidence mounts, and now the common wisdom is that IBM picked the "wrong" CPU for the "right reasons".

And then you still hear lots of things that don't match reality. Or, at least, I still hear a lot of things that don't match what I know and remember on the subject. That's part of the reason I wrote this rant, to tell what I remember of things.

But you never hear what I think is the real reason, and that's the real reason I wrote this.

The short version of what I understood at the time was the reason -- and I have seen no real evidence to the contrary -- is this:

(1) IBM didn't choose Motorola's 6809 because Motorola did not support it like they should have. Motorola didn't support the 6809 like they should have because they were afraid of it eating into the 68000's market.

(2) IBM didn't choose Motorola's 68000 because they (IBM) were afraid that a personal computer based on the 68000 would eat into their market for their mid-range (minicomputer-class) computers based on the System 3.

That's the conclusion, and the rest of this rant is about where that conclusion comes from.

]

You need to understand the situation in 1979-81 clearly.

You have to remember that IBM was not officially considering entering the personal/home computer manufacturing business. 

In another company, the project to develop the PC might have been called a skunkworks project. But the PC project had even less official recognition. Yes, they were working separately from the main company, yes, management kept hands-off. (Some had even washed their hands of it.) Work was performed in secrecy, and it was not only started without contract or official directive, it was mostly complete by the time upper-level management acknowledged it.

IBM already had their blue-sky research projects, which was something more akin to the Skunk Works at Lockheed. This was different.

Not to say that Skunk Works and blue-sky were completely free from adversarial management, but the PC project was pursued in a much more adversarial management environment. The engineers who built the initial prototype were permitted to do so by their manager, who acted in specific contradiction to direction from the next-up level of management. 

At least, that's the story I heard several times while working internships for IBM, and those stories matched what I was seeing, where later stories do not.

Of course, those who let the project move ahead were, to more-or-less extent, putting their own jobs on the line for the results.

IBM's marketing and engineering did not want to deal with the threat of microprocessors in general-purpose computing devices. The attitude I heard was, of course microprocessors can't do the job. They are strictly for controls devices and calculators. 

And it wasn't exactly a mistaken attitude. All existing microprocessors at the time were missing elements that were important to general-purpose computing -- memory management hardware, direct memory access input-output channels, a hardware timer for dividing the CPU's time between tasks, proper hardware division of task resources, .... And the list goes on. 

Microprocessors are still missing much of that list. But even the "big" computers of the time didn't have all these things in place, either. So it was, in fact, ignoring reality.

I will mention this attitude again further down, but this is enough to get a feel of how things were at IBM.

Apple and Commodore's history is pretty well known. That is to say, I was not close enough to them to add much, so I won't. Everyone knows that Apple IIs were selling well in business and education markets, and Commodore's offerings were just behind in business and education, but were ahead in personal/home use sales, and eating away at the dedicated games machine market.

But Radio Shack's history is not so well known. This should not be surprising. Radio Shack really didn't have an approach to write about. The TRS-80 sort of fell into their lap. 

(Again, I was listening to local management discuss things while it was happening. I was trying to be a Radio Shack salesman in Odessa, Texas when the first TRS-80s were delivered. I got to unpack our demo unit and write something up as a demonstration program.) 

The guy who designed the Z-80-based TRS-80 original model (now called Model 1) just kludged some stuff together from demonstration circuits published by Zilog and hobby circuits from the hobby industry. I saw the circuit diagrams, and I knew enough to tell where some of the short-cuts had been taken that were a bit beyond specifications for the parts. It was intended as a proof-of-concept, but Radio Shack had no engineers at the time to actually fix the design, and had no motivation to do so, either. It was part of their culture, and they were looking to compete with folks like Tramiel.

Now, I took a break from the technical world during 1979 and 1980, to serve God according to my understandings and belief. When I returned, Radio Shack had hired a few engineers and was half-heartedly trying to clean up the Z-80-based model designs. 

And they had had another personal/home computer design fall into their lap, the M6809-based Color Computer. This one was a Motorola example circuit, with just a little modification. 

The 6809 was potentially more powerful than the Z-80, powerful enough to go head-to-head with the 8088 in certain applications, powerful enough to build a minicomputer-class microcomputer with. But Radio Shack didn't have the engineers or the motivation to do anything but sell the thing as a game machine -- or as a toy for geeks. 

Microware had been able to get their OS-9/6809 operating system running on the Color Computer, and it was suddenly a serious business and/or industrial controller class machine -- a seriously (woefully) hobbled business machine, but squarely in the same class as the IBM-PC would shortly be released for. So Radio Shack let some engineers cobble together a floppy controller that could support OS-9/6809 and could be plugged into the game cartridge port -- instead of investing in designing a true business-class machine as an upgrade to the Color Computer.

Admittedly, Motorola was not really supporting the 6809. They were busy selling into the microcontroller market as many true 8-bit 6805s and mixed 8/16-bit 6801s as they could manufacture, and they considered their future to be riding on such smaller microcontrollers and on the 68000.

We should remember that the nascent PC market was not nearly as important to Motorola (or to Zilog) as it was to Intel or Commodore. Motorola sold several orders of magnitude more microcontrollers than any CPU manufacture has ever sold microprocessors for personal computers. And it was hard to argue with their attitude. It would be fifteen years down the road before mind-share issues would be eroding their position in the controls market enough to get the full attention of Motorola's management.

Focusing on controls was not a mistake for Motorola. Not seeing the mind-share problem was. And only a very few engineers anywhere were seeing the potential uses for PCs as communications devices at the time, so it's not too surprising that Motorola didn't recognize why the PC was important to their future.

That was how things stood when I returned in late 1980, intending to work my way through school as an electronics tech.

So, Radio Shack needed to have better engineering to do anything with the better CPU technology that had fallen into their lap -- twice. 

Commodore's Tramiel wasn't the only one in the industry who thought engineering should be sacrificed for price and immediate profits, and was not the only one who found himself left to drink from the marketing stream he himself had polluted.

So what about engineering? Was it really so obviously necessary back then as it seems in hind-sight?

Motorola was fighting with engineering missteps in the 68000. Intel was struggling with over-engineering the iAPX 432 (and under-engineering the 8086, but ...). Zilog was struggling with similar problems with their Z8000. 

The first 32-bit addressing CPUs (microprocessors or not) all bore the marks of engineering specs that were way too ambitious in application areas that were poorly understood -- too much engineering without foundation in real-world experience. (We're still doing it.)

Consider this -- the 32-bit CPUs existing before the 68000 were not microprocessors, of course. They were all leading-edge hardware in mainframes, and the companies that produced them were protecting their inventions and technology with strict secrecy. 

Large memory systems, sharing a processor between multiple tasks, and coordinating the work of multiple processors were all brand new application fields for most of the industry. 

Real technology was just not available. That is, what was available other than theoretical technology was really hard to come by. And when you engineer things that complex with little real-world experience to guide you, you're going to make mistakes.

Insane competition drove the industry to push ahead into the 32-bit world a lot harder and faster than was safe or wise. In some senses, IBM's and Motorola's hesitation made sense.

Usually, you hear that IBM's options to the 8086 were Motorola's 68000 and Zilog's Z8000. Of the three, the obvious choice was the 68000, and the usual question is why it was not chosen.

Remember, nobody knew what they were doing. 

The 68K was the best of the actually available options, but Motorola's design for handling the big memory space was too ambitious, relying on false understandings of the underlying problems. That was where the "bugs" were, although nobody really had a better handle on it at the time. 

Three specific design misfeatures in the 68000:

(1) Complexity, and the expectations that induced, although that didn't get really out of hand until the 68020. 

Nowadays, the 68000 seems relatively tame, but it was at least an order of magnitude more complex than, say, the 6809, and we lacked testing and design tools for the complexity then. 

You could use 8-bit engineering on the 68000, but you had to ignore all the cool features of the CPU to do so. That was hard for engineers designing for the 68000 to stomach, not to mention for management and marketing to justify. 

I think it was less the cost of supporting it all than having to plan on starting with a machine that both you and the customer would know was going to be replaced by a more complete re-design within the year.

So, why not build the more complex design to start with? That was what a lot of tech companies tried to do, and what a lot of them failed at.

[EDIT 20240422: 

IBM themselves had a department trying to build a real computer based on the 68000 at the time, the IBM Scientific 9000 or 9000 Scientific, depending on what docs you were reading.

]

It took too long to test, especially to the expectation levels we had then. We were too afraid of bugs in non-critical software and hardware. We were missing tools, and management was too scared of sinking money into the projects to fund developing them.

We are now used to the market itself being a required testing stage, but we weren't used to that idea back then. (Think about how your "smart" phone does so many things you don't want it to now. That's misfeatures, you know -- bugs -- that you are testing.)

For the record, note that the 8086 was about half an order of complexity more complex than the 6809, although it was less complex than the 68000.

(2) The original 68000 was not directly fully 32-bit addressing when using position-independent code. (Absolute addressing code, yes. Position independent, almost, but requiring software shims for modules larger than 64K.) There were missing 32-bit constant offsets in certain addressing modes. 

Position independence was a great plan, and Motorola should be commended for embracing it. But by only giving offsets of +/- 32K in those indexing modes, Motorola had built the 68K with a hidden barrier to overcoming the 64K module size barrier when designing for position independent code. You had to use two instructions and a register (a 32-bit load immediate in addition to the register-offset indexing mode of the instruction you wanted) to get full 32-bit range with position independence. 

This is most of the reason Codewarrior for Mac had a small memory model similar to the x86's small memory model. If the 68000 had had full-range constant offsets, the memory models for the 68000 could have been blended, and programmer wouldn't have had to plan for it, and the small model would have disappeared as the compiler matured.

Imagine telling your manager about that, after you successfully lobbied for the 68000 instead of the cheaper, but less capable 8086, because the 8086 had the barrier and the 68000 didn't. (Then think about trying to explain why position independence, which wasn't really achievable on the 8086, was important.)

Well, two instructions and one data register on a CPU with 8 data registers is not nearly the impediment that four-to-eight instructions and the only accumulator the CPU has on a CPU with only one accumulator. The small and big models on the 68000 could have been blended anyway, but we (the market) had this antipathy to compilers that took more than two optimizing passes and then added another optimizing pass in the link phase. 

(It would be a few years before we as an industry generally recognized that optimizing too early was a big problem, but that's a rant for another day.)

(3) Motorola made some problematic design decisions in how the 68000 handled exceptions in the intransigent case of memory page faults and such. As a result, there was not enough information about the cause of the exception to recover and continue, which made it difficult to share memory between running applications safely.

Intel just didn't handle it at all in the 8086, which left them able to quickly recover from the mistakes of trying to implement too much complexity in the 80286 when they moved to the 80386. Most companies who needed to deal with memory exceptions for the 80286 said they would try to implement the exception stuff in their next software upgrades, but by the time they were getting started on the next upgrades, Intel had the design for the 386 ready. The 80386's approach was much simpler, and nobody needed to go down the 80286 rabbit hole any further.

Motorola tried to handle those exceptions in the original 68000, but got it slightly wrong, and that, more than the extra clock cycles, was what kept the 68451 from being the MMU in any of the major workstations built on the 68K. Engineers understood the problems of wait states, and expected that the newer versions would be able to run with fewer wait states. They could expect wait states to be handled in hardware. But the exception handling misfeature meant having to plan on re-writing the very code that they thought they only wanted to write once. 

Motorola did fix that in the 68010, which was released to the market in 1982. Unfortunately, they did not fix the 32-bit constant offset problems until the 68020.

Now, it should be noted that, ultimately, problems in the code for handling memory management is still biting us in the form of microcode vulnerabilities, and somebody has to regularly rewrite it. (Remember the Heartbleed vulnerabilities, for example?) ARM64 suffers less than AMD/Intel64, but the CPU vendors are still struggling with it. It's a very difficult problem to solve well.

Which means that the mistakes in the 68000's exception handling were not really a 68000 problem,  they were a general problem for the whole industry. But everyone still likes to call it a 68000 problem, because no one really wants to admit we as a race don't already know everything we need to know for handling today's problems yesterday.

Now, Intel did have an advantage of sort-of learning from their misadventures with the iAPX 432. That is, they just decided to punt on the problem with the 8086, and gave it non-enforcing segment registers -- which aren't really segment registers. It was a (not very good) non-solution with interesting and sort-of useful side-effects that cause more problems down the road.

The 8086 segment registers, instead of being the width of the 20-bit addressing of the 8086, were just sixteen bits slid over four, which made the segment registers clumsy to work with as base registers for large arrays and such. You ended up needing four to eight instructions and two or three registers to properly handle 20-bit pointers.

And there were no segment limit registers, which is why I say they were not really segment registers. Implementing segment limits in software on the 8086 essentially doubled the already excessive instruction overburden.

In a sense, the pseudo-segmentation was used like a cheap alternative to bank-switching, avoiding the use of external hardware to achieve expanded memory range. Ironically, though, until the 80386 was available, external hardware bank switching was still used in addition to those pseudo-segment registers on the 80x86-based PCs with large memory requirements.

The segment registers did not really solve the underlying problems, nor did they contribute to a proper external hardware solution, they just allowed (with the cooperation of the market) the problems to be swept under the rug until the 80386 was available.

Yes, as I mentioned above, the 68000 had sort-of similar problems, but not nearly to the same degree. If you needed segmentation on the 68000, it was just a matter of index modes, and even with the shortcoming with constant offsets, it only took two instructions and a register to take care of, not four-to-eight instructions. (You just had to ignore the obvious CHK instruction if your segments were going to be larger than 64K -- until the 68020, when that was fixed. But, back then, nobody wanted to waste time bounds-checking things anyway.)

(4) (Note, this is more than three.) The 8088's 8-bit external was cheaper! -- but no, not really. 

You may have heard such nonsense about the 8-bit external 8088 being cheaper to design for than the 16-bit external 68000. Let's calculate this:

-- Yes, the entry level model would have had sixteen 16Kx1 dRAMS instead of just eight. 

No, that would not have broken the bank. The initial sticker price could have added the extra USD 32.00 or so without breaking any significant price barrier. (Go look up the original prices.) That differential would drop as production ramped up.

The RAM configuration would have been 16 chips wide instead of 8, but that was only a small routing problem -- eight extra wires. The max on-board RAM for the mainboard in the original version could have remained at 64K, although in configuration of 32Kx16 bits (32Kx2 bytes).

-- Yes, the expansion bus did present a problem. Catering to the 8088's limitations allowed sweeping the expansion bus width under the rug for a couple of years, limiting performance options that could otherwise have been had if the 8086 had been a planned option from the beginning. 

Wider connectors were more expensive, but only until ordered in large numbers. Total added to the initial price? Maybe USD 5.00. And that would disappear quickly as production ramped up.

Really, the narrower 8-bit expansion bus was solving a different problem than cost.

-- The 8 kilobytes of BIOS ROM was done as 4 2Kx8 ROMS, and the only difference necessary would have been the 16-bit wide data bus configuration. 16-bit data meant pairing the ROMs, but 8 kilobytes is 8 kilobytes whether done as 4 unpaired ROMs or 2 pairs of the same size ROMs:

4x(2Kx8)== 8Kx8 bits => eight kilobytes

vs.

2x2x(2Kx8) == 4Kx16 bits => eight kilobytes

Yes, the engineers would have had to take a little time learning the 68000's superior instruction set to be sure they weren't wasting space and cycles in the 4 2Kx8 BIOS ROMs, whereas they were already familiar with the 8086's instruction set. 

Marketing types with no patience to understand what they were talking about tended to say things like, 

But the instructions are 16 bits wide! That's going to be twice the memory to store programs!!!!!!

That is potentially a legitimate concern with 32-bit RISC instruction sets, which are usually not as densely encoded as either the 8086 or 68000. 

But instructions in the 8086 are variable width, in 8-bit chunks. Instructions in the 68000 are also variable width, in 16-bit chunks that do more work than the 8-bit chunks of the 8086. Motorola was careful there.

(Just for the record, neither the 8086 nor the 68000 is as densely encoded as the 6809, but if the 6809 instruction set is expanded to handle addresses larger than 16-bits, you lose some of that density.)

-- What else? Peripheral chips? 

The 68000 had instructions specifically to enable using 8-bit wide peripheral parts without having to adapt the peripheral's 8-bit wide interface to the 68000's 16-bit wide data bus. It might have felt a little "awkward" or "non-ideal" to some engineers, but most experienced engineers would not have even blinked an eye. 

As a specific case-in-point, the original IBM PC used Motorola's 6845 to generate video. Motorola had a reference design for using that exact chip with the 68000 (and, indeed, hardware for the 68000 based off that reference design were not unknown in the industry).

(5) Lack of software isn't exactly a misfeature, but it is often invoked as a reason the 68000 wasn't ready.

Remember that CPM/86 had not entered into development when the IBM PC unofficial project began. Whatever CPU they chose, they were either going to be dependent on the CPU's manufacturer for an existing OS, developing their own, or getting a third party to develop one for them. Choosing the 8086, they initially turned to Digital Research. 

Note that it was already known that 8080 or Z-80 software needed more than just re-assembling the source code with an 8086 assembler. Transliteration was possible to an extent, but cleaning up the transliteration did take time.

Considering the amount of Z-80 software that was quickly (and crudely) transliterated to the 68000, asking Digital Research to develop a version for the 68000 would not have been unreasonable. Likewise, they might have reached out to Technical Systems Consultants or Microware for a 68000 version of their OS products for the 6800 and 6809.

At this point, you should be able to see that all of the usually mentioned strikes against the 68000 were not strikes at all. Balls. The 68000 really should have walked the bases, so to speak.

Setting the 68000 aside for a moment, I've mentioned the 6809 a bit above. You may be aware that Apple considered the 6809 for the Macintosh, and even wired up a prototype before deciding that the Macintosh really needed more address space.

I've heard, but have not corroborated, that the 6809 was also considered by IBM for their PC somewhere along the line. Would it have been a bad choice?

On the plus side:

(1) The 6809 was designed from the beginning to handle high-level languages, multi-tasking and such. I mentioned Microware's OS-9/6809 above. Uniflex from Technical System Consultants was also available in 1979. (TSC's Flex for 6809 was available almost as soon as the CPU was.) IBM would have had two relatively mature quality OSses and a good developers' environment and community from the outset, and much less aggressive partners to work with.

(2) Motorola did have a page-mode MMU part for the 6809, the 6829. External MMUs do tend to slow a processor down a little, but, with the 6829, the 6809 was able to compete with minicomputers -- minicomputers from the mid-1970s, but that's not bad considering that those mid-1970s minicomputers were still very actively used into the mid-1980s.

(3) Motorola had a floating-point ROM for the 6809 as well, the 6839. It was not as fast as floating point in hardware, but it was ready, and cheaper.

(4) The 6809 was fairly well-known at the time for its ability to handle graphics.

On the minus side:

(1) The 6809 did not have segment registers. Breaking the 64K barrier would have required bank switching or full paging, and doing large arrays with bank switching or full paging requires a bit more code than even with the 8086's sloppy segments.

(2) The 6809 did not have hardware divide. And hardware multiply was only eight bits wide, so you had to put four of those together to get a 16-bit multiply. That also slowed down the floating-point operations.

(4) Motorola was not talking about extending the design. They were focused on their profitable 6805 and 6801 CPUs.

That last was possibly the killer for the 6809.

If Motorola had been showing signs of actually incorporating the 6809 as a core in microcontrollers similar to the 6801 and 6805, IBM could have been confidant in being able to get them to build a 6809 with something like the 6829 memory management unit built-in, and integrated MMUs tend not to slow CPUs down nearly as much as external MMUs. 

And they could have had reason to expect Motorola to extend the 6809 architecture. Simply adding linear segment registers and full 16-bit hardware multiply and divide to the 6809 would have made it head-to-head competive with the 8086, in spite of the 6809's 8-bit architecture. Extending the architecture in something the way the 8086 was extended to the 80386 would have been no particular problem, either.

But such a 6809 would also, as the theory goes, have eaten into the 68000's market. 

(It is true that adding segment registers would have conflicted with the page-mode MMU operation that had become sort-of expected in 6809 software, requiring a small bit of work to bring, for example, OS-9/6809 up in a segmented model instead of a paged model.)

I personally think Motorola should have given the 6809 more attention, as I have mentioned elsewhere in this blog.

Anyway.

No, not anyway.

The Z8000 was not ready, and was not mature.

The 6809 was ready, and was mature. It would have made a very good base for a PC, if Motorola had had plans to upgrade the CPU architecture in future offerings. Such plans were beginning to be obviously not in evidence.

The 68000 was ready, and Motorola was clearly committing to supporting and upgrading the architecture. Anyone who tells you otherwise either does not know or is ignoring the overall situation in the marketplace.

The 68000 was also 8-bit capable, contrary to what some have said.

Even if it would have taken an additional six months to a year to qualify, IBM never did a proper qualification of the 8088 PC design either, and there was no formal marketing study that defined a market window to be met or anything like that. And there really was no reason to expect it to take any longer.

It was not theoretical delays, not deficiencies in the 68000, per se, not anything like that.

So, drum roll:

Here is the real reason, from what I heard at the time, and I still believe it:

The 68000 was powerful enough to allow building a microcomputer that would have competed with IBM's system 32/34/36/38 series of minicomputers. 

(Yeah. I saw that, when I was doing the internship with IBM. The system 3 CPU even looked a little like the 68000, superficially, inside. We could guess they'd have done better by moving the System/3 series to their own custom version of 68K, but thinking about that requires enough hypothesis contrary to fact to push us into the realm of writing alternative reality SF

And we can give a nod to IBM's first exercise in putting the system 360 on a microprocessor, while we're at it. Look up the IBM Personal Computer XT/370. That was real history.)

The x86 was not powerful enough for that. That left IBM's marketing team able to imagine they had breathing room.

The real reason was the same reason Motorola didn't support the 6809 the way they should have:

There was this meme that seemed to go around every marketing department back then --  

"MUST. NOT. COMPETE. WITH. OURSELVES!!!!!!!!!"

(I think we all now know that, when you're careful to avoid going to all-out war with yourself, competing with yourself keeps you from getting complacent.)

So the IBM PC project had to be kept hidden. And that is why it resulted in a non-optimal product. And the casual engineering of the x86 PC itself ultimately almost took IBM down with it, anyway. 

Microsoft and Intel would avoid the back-swill of their sins with a lot of planned smoke-and-mirrors, always (barely) able to make it look like they were leading the way out of the mess they themselves were creating. 

And that is probably as close as you can get to the real story of why bad engineering prevailed, but was only truly successful for Microsoft and Intel, both of whom are now eating and drinking the pollution they made -- along with the rest of us.

The real question is

Why did such a non-optimal design succeed?

Full answer takes us deep into religion. I'm not going to ask you to go there with me today. Maybe some other day. 

Quicker answer, if only partial: 

The (business) world wanted spreadsheets along with the word processors. We didn't all know exactly what they were, but we wanted a bigger and better calculator that would allow us to do with our accounting books what word processors allowed with prose records. And we wanted it cheap, and we wanted it big. And we wanted it on our desks.

The Apple II got us close to giving us that, but spreadsheets on the Apple were not as intuitive as word processors were becoming, and were somewhat limited in size.

Both the x86 or the 68K supported more intuitive spreadsheet apps capable of handling larger spreadsheets, but the 68K was a threat to marketing departments.

That was why the IBM PC snowballed. Small, weak things are sometimes bigger and badder than big strong things. (That's the short version of the religious discussion, too, by the way.)

There really was only one way to have avoided the mess, and that was for IBM to have resisted the temptation to try to jump into the front of the race. (And I'm busy writing a not-very-good novel about that, when I'm not too tired after a day of delivering mail to keep food on the table, so I'll forego talking about that here.)

Monday, August 9, 2021

Guessing Which Motorola Microcontroller Part It Is (6801/6805/68HC11/6809)

Wasted too much time on this. 

This is extracted from my response to a post to the Facebook vintage {Computers | Microprocessors | Microcontrollers} group, asking for help identifying a Motorolo-logo microcontroller found in a washing machine with an apparent custom SOC part number ZC85148L, with an apparent date stamp from early 1984. 

There were many guesses as to what the ZC85148 was, and I thought I'd put my guesses and reasoning out here, to make them more available for searching:

I'm guessing either 6805/68HC05 or 6801/68HC01.

6805 was essentially a stripped-down 6800, with only one eight-bit accumulator (A), one eight-bit index (X, yes, eight-bit), bit instructions, expanded indexed modes, better power-saving stuff, lots of timers, some analog-to-digital, and other integrated I/O to choose from. At least some parts included hardware 8-bit multiply.

6801/68HC01 was exactly the 6800 with a few new 16-bit instructions for the double accumulator A:B pair, better X handling, hardware 8-bit multiply, and better power-saving mode stuff.

My reasoning is as follows:

If the chip were a bare CPU, it would need separate ROM and RAM parts on the circuit board, and such were nowhere in evidence. From the late 1970s until Motorola spun off the microprocessors business and renamed it Freescale, they provided Systems-On-a-Chip (SOC) semi-custom microcontrollers which included RAM, ROM, and I/O on chip. It's a pretty safe bet that the part was an SOC microcontroller.

I am not aware of any 6809 SOC products from Motorola, so, since the circuit board showed no evidence of ROM or RAM, I'm pretty sure that it is not a 6809 variant of any sort. (If there were any special-order 6809 SOC microcontrollers, that would be interesting to hear about.) 

The 68HC11 did start shipping in 1984, so it could possibly be a 68HC11. 

(68HC11 is a 6801 in HCMOS, with an additional Y index register and a pre-byte that converts X-indexed op-codes to Y-indexed opcodes, plus hardware integer and fraction divide, bit instructions, and a little bit more. Some people confuse the 68HC11 with the 6809. I discussed the differences between the 68CH11 and the 6809, comparing their architecture and lineage, in another rant here, several years back: https://defining-computers.blogspot.com/2018/12/68hc11-is-not-modified-6809-and-what-if.html. There are a few errors there I need to go back and correct sometime, but they are not errors of substance, I think.)

I don't have solid information on when the first HC08 microcontrollers started shipping, but my impression is no sooner than the late 1980s. So  I'm guessing it was not an HC08 or HCS08 or any of the later extensions thereof. 

(HC08s were 68HC05s with a high-byte extension for the X register and some other useful stuff, including hardware 8-bit multiply and divide.)

Other possibilities -- I understand that Motorola did second-source at least one other company's CPU in the 1970s. Whether the Intel 8501 might have been one of those, well, it seems ludicrous, but I have some conflicting memories. I think they had mostly gotten out of that business by the mid-1980s

It is my memory that they dabbled in manufacturing IBM compatible desktop PCs in the mid-to-late 1980s, but I don't remember whether that included manufacturing their own 8086 compatible CPUs, or second-sourcing Intel's. Anyway, that ended up only for desktop, and management quickly recognized that business model was not going to be profitable for them (and had been a marketing misstep).

I did also see announcements and engineering materials for 6502 core SOCs from Motorola, somewhere around 1986 or '87, IIRC, but those also seemed to have been dropped pretty quickly. It would not have been there in 1984.

And it is also my understanding that Motorola provided manufacturing for some mil-std microcontrollers, but I think that was only to the military and maybe NASA. Those were 16-bit, and not based on any of the 68XX or 68XXX series. I suppose it would not be impossible to see something like that in a washing machine controller, but it would be overkill.

There were also 4-bit and 1-bit SOC microcontrollers that Motorola produced in the late 1970s, but I don't remember any of them in 40-pin dip. I think they would be a bit underpowered.

I don't think the 88000 RISC series was even announced yet in 1984, and the Power architecture discussions with IBM and Apple had not even been imagined yet. There was the one-off custom implementation of the 360 architecture borrowing from the 68000. We can be sure that none of these would have been in a 40-pin chip in a washer in 1984.

Which is why my guesses come down to the 6805/68HC05 or the 6801/68HC01.


Monday, March 1, 2021

What Is (the Programming Language) Forth?

[EDIT 20240526, add-overview: ]

This needs a higher-level introduction, I think.

Forth is a simple and compact combination of minimal BIOS and library, rudimentary OS, and a very flexible command-line programming language shell that encourages modifying the Forth itself as an approach to developing applications. 

It's compact and simple enough that, using hardware that was common back in the 1970s and '80s, a single moderately competent engineer could put together a running system that can host its own development from scratch in a few months.

Because it is so easy to put a system together, and is so flexible, Forth is a bit of a siren. Many incompatible versions have been developed, and many engineers have found themselves starting grand plans that seem doable with Forth, but end up foundering on some of the gotchas that the simplicity hides.

The key features, I believe, are 

[EDIT 20240526, end add-overview; ]

  • (1) The colon definition grammar.

  • (2) The post-fix expression grammar.

  • (3) Two stacks.

  • (4) On-line dictionary (symbol table) at development time, and, unless explicitly removed, at run-time.

  • (5) The ability to blend run-time with compile-time at development time.

But this is way too loose. It leaves out all sorts of implementation details that non-trivial applications depend on.

[EDIT 20240526, expand-on-key-features: ]

I've left out of the above one key feature that many fans of Forth seem to think is essential, something called 

  • (6) an inner interpreter, or a virtual machine.

And I think I have good reason, but it does require adding a feature that is considered optional by most, but I need to expand on the above -- and add a few more key features -- before I pick that question up.

Deliberately taking up the above key features in a different order -- 

  • (3) Two stacks is an optimization for parameter passing.

In typical compiling convention, the call stack interleaves parameters and call-return information on a single stack. It's practically an assumed convention, so much so that many central processing units directly support the creation of the call record in a combined stack frame. It's also very inflexible and time-consuming, leading to significant complexities in interpreters, compilers, and the run-time models they work under and support.

One particular complexity for compilers is compiling procedure/function calls as in-line code instead of called code, requiring complex evaluation of the procedure or function, and adding significant bulk to the output code wherever the calls are in-lined.

The combined stack is also more fragile, making it much easier to accidentally or deliberately overwrite the call-return information, resulting in application behavior outside what the application engineers designed -- in other words, (additional) bugs and vulnerabilities.

The only advantage of the single-stack approach that I know of is eliminating a memory management segment. Memory management is perhaps the most complex common problem in general applications, and reducing the segment count from, say, four to three, is generally viewed as vital.

By splitting parameters from the call-return information, calls can be streamlined to the point that in-lining makes sense only for the very most simple functions. This will significantly reduce the complexity of the compiler, as well.

As a bonus, keeping the call-return in a separate segment makes it significantly harder to overwrite, eliminating one of the most common source of bugs and vulnerabilities.

  • (2) The (mostly) post-fix expression grammar kind-of falls naturally out of the split stack calling convention.

Post-fix is not required, and in-fix and pre-fix expressions are also simplified by the split stack convention, but a primarily post-fix expression grammar allows all elements of the language to be implemented simply as if they were called functions with results that are not limited to scalar results. 

Post-fix is really simple with the split stack, again, reducing the footprint of the interpreter/compiler.

I should probably note here that most kernel Forth interpreters do not check parameters on parse or call, which is a convenience when working at the low level, but also a vulnerability. 

One thing needed for Forth to be more accepted in the industry is a higher-level compiler that checks parameters at compile-time (unless directed not to). The lower-level compiler can be kept for debugging-level work and such, but a higher-level compiler is necessary.

  • (1) Colon (and other) definition grammar is an exception to the general post-fix expression grammar. 

If implementing a stripped-down Forth, the colon definition is the only one necessary, but constant, variable, vocabulary, string, and compiler-compiler (etc.) definition grammars are a necessary convenience if you're not aiming for the most stripped-down implementation possible. 

These defining grammars are (usually) all pre-fix or in-fix grammars. They can be done post-fix, but the grammar gets really convoluted when you do that. 

Colon definition reads backwards from the usual English usage of the colon, but it gives one the sense of defining functions, procedures, variables, constants, vocabularies, compiler-compilers, etc. as words with definitions.

  • (4) The symbol table is implemented as a dictionary, or, more accurately, as a collection of vocabularies, and the interpreter/compiler is implemented as the initial dictionary.

This is what makes the interpreter/compiler facilities available to the programmer at development time, and even to the end-user at application run-time, if the application is so written.

Essentially, this dictionary implements what are called libraries in other languages, but it makes the whole language available to those who dare use it. 

  • (5) And, as I mentioned just above, the dictionary can remain accessible to the end-user, blending development-time with compile-time and run-time.
Often, in definition grammars, common patterns call for delayed binding, but I have only once seen a Forth that tried to implement delayed binding. It gets really trippy, and really is (usually) not necessary -- because the compiler itself remains available to the run-time unless you strip it out.

Alarm bells are ringing in system engineer's brains when they read the above, but without delayed binding, access to the compiler can be controlled by the application.

And that brings us to 

  • (6) The inner interpreter provides the framework within which the parts of the run-time not provided directly by the CPU are implemented. 

This was useful back when the concept of a run-time architecture was not well understood by systems and software engineers, and even by CPU architects. It is completely superfluous with CPUs such as the 6809 and the 68000, and even mostly superfluous for the x86 and such.

An inner interpreter can be useful in defining a debugger without having the debugger know all about every feature of the CPU, as well, since all code ends up back at the inner interpreter, and all the core features are defined in the initial dictionary.

But that debugger can become cumbersome and hard to control when investigating bugs in the interpreter itself, or bugs in the programmer's understanding of the runtime and interpreter.

Moreover, the inner interpreter usually gets in the way of combining Forth definitions with the libraries of other languages. For instance, combining Forth definitions with C libraries will require implementing a C compiler and libraries that understand and run within the runtime defined by the inner interpreter.

Engineers who write Forth code tend to become impatient with the idea of implementing an entire C compiler in Forth (which is the only real reason we don't see that happen, but is sufficient reason).

For this reason, I have been and am inclined to try to implement the Forth virtual machine run-time as if it were an extension to the CPU rather than a virtual machine defined by an inner interpreter.

This will mean native CPU calls, which will mean that the compiler has to know how to compile at least the native CPU call. And it will mean that the minimal kernel must have some minimal debugger features and know at least how to compile and see the native call instruction. Yeah, it's not as easy, but having now done several fig-Forth implementations (including a buggy conversion of the 6800/6801 fig Forth to the 6809 (osdn can take a while to come up) and a fairly bug-free conversion of the same to the 68000) I'm even more inclined to think it's worth the trade-offs.

Returning to the fact that Forths are easy to roll your own, and easy to make non-standard, ...

[EDIT 20240526, end expand-on-key-features; ]

Comparing it to the programming language C, it's like saying that, not just all K&R C compilers are included as C, but all the different versions of Small C and Tiny C. (And  maybe even Objective C, Javascript, Java, Ruby, PHP, and a number of other languages that borrow heavily from C syntax and grammar?)

The solution?

fig-Forth is one group of dialects that have a lot in common. We could define and develop a standard fig-Forth.

I've transcribed the 6800 fig-Forth model and optimized it for the 6801. In addition to some I/O bugs, there are enough differences from the 6502 model to cause problems for non-trivial applications.

I have a near-fig-Forth I call BIF-6809 (and a non-functional one I call BIF-C). They use a binary tree symbol table, and that alters the language enough to make it unreasonable not to give them separate names. (Double negative intentional. Not just reasonable to have separate names, but unreasonable not to.)

Forth77 and Forth83 are separate languages, and should be treated as such. And they should be referred to by their complete names.

SwiftForth's language should be referred to as SwiftForth.

ANSI Forth should be renamed. Call it CommitteeForth or something.

The name Forth, unadorned, should be reserved to whatever Charles H. Moore (the original author of Forth) wants to call Forth.

(I should note that Moore himself calls his own dialect ColorForth. He's leading out, here.)

We need to make the nomenclature a part of the dictionary/symbol table. A word called version should bring up a version number, sure. 

We need a word that returns (without printing it to the terminal) a string containing the name of the language/dialect. Maybe even one word for the language family name and one for the dialect. That would give us a base point, after which it would become possible to check what kind of glue needs to be brought it, to make a particular source code compilable with a particular compiler.

After some consideration, I'll suggest the following four new words:

* language ( --- adr )

Returns a string containing the language name, "Forth". (This would allow distinction from derived but different languages.)

* dialect ( --- adr )

Returns a string containing a dialect name, "fig-Forth", "Forth77", "Forth83", "SwiftForth", "gforth", "BIF-6809", etc.

 * sub-dialect ( --- adr )

Returns a string containing a modifier of the dialect name. 

* target-cpu ( --- adr )

Returns a string containing the CPU targeted, such as "6502", "6801", "Z80", etc.

In addition,

* version ( --- ud )

Should return an unsigned double integer in which the first byte contains the major version number, the second byte the minor, and the third and fourth contain a sequencing number within the minor version.

To further aid tuning source to the host language, run-time, CPU, etc., dialects which adopt this practice should also adapt the practice of defining words that describe such things as cell width, sign representation, defined boolean constant to set a logical true, etc. We could take some inspiration from C's (original, sparse) limits.h include file for this, but it should not duplicate the contents of limits.h.


Thursday, November 19, 2020

Only install system updates from within the OS – from the settings menu.

Got an update notice from Playstore. 

I went to System Settings to update.

Nothing.

That means Playstore apps can send notices that look like system update notices.

Fake notices.

From apps that probably install malware of some sort. I need to tell Google about this but I don't have time until lunch.

Lesson?

Only install system updates from within the OS – from the settings menu.

Monday, October 26, 2020

6800 Example VM with Synthetic Parameter Stack

In Forth Threading (Is Not Process Threads), I commented on the fig Forth models insisting on using the CPU stack register for parameters. The code I presented for the 6800 follows that design choice, but the code I presented there for the  6809 and 68000 does not. Here I'll show, for your consideration, the snippets for indirect threading on the 6800 done right (according to me) -- with the CPU stack implementing RP and the synthesized stack implementing Forth's SP. 

You'll note that the code here will end up maybe 5% slower and 5% bigger overall than the code that the fig model for the 6800 shows. But it will not mix return addresses and parameters on the same stack, and it should end up a bit easier to read and analyze, meaning that it should be more amenable to optimization techniques, when generating machine code instead of i-code lists.

In the code below, the virtual registers are allocated as follows:

  • W    RMB    2    the instruction register points to 6800 code
  • IP    RMB    2    the instruction pointer points to pointer to 6800 code
  • * RP is S
  • FSP    RMB    2    Forth SP, the parameter stack pointer
  • UP    RMB    2    the pointer to base of current user's 'USER' table  ( altered during multi-tasking ) 
  • N    RMB    10    used as scratch by various routines

The symbol table entry remains the same:

    HEADER
LABEL CODE_POINTER
    DEFINITION
So we won't discuss it here. (See the discussion in the post on threading techniques.)

I won't repeat the fig model here. I'll be bringing more of it into the discussion than I did in the post on threading techniques, refer to the model in my repository, here: https://sourceforge.net/p/asm68c/code/ci/master/tree/fig-forth/fig-forth.68c.

This is an exercise in refraining from early optimization, so the parameter stack pointer will point to the last item pushed, even though it's tempting to cater to the lack of negative index offset in the 6800 and point to the first free cell beyond top of stack.

The inner interpreter of the virtual machine looks the same, but the code surrounding it changes. 

PULABX is only used by "!" (STORE), so let's start the rework there:

Forth Threading (Is Not Process Threads)

Code-threading in the context of Forth deserves its own rant to help untangle the knot I picked up in Computer Languages -- Interpreted vs. Compiled Forth?. This is not a complete treatment, only background and an overview. YMMV.

Forth Threading (Is Not Process Threads)

If you've worked much with low-level code, when I mentioned (in comparing compiled and interpreted Forth) that Forth compiles definitions to a list of the addresses of definitions, you will have recognized that there is a recursion there that has to be broken at leaf calls if a Forth VM ever wants to get any useful work done. 

At some point, real code has to be executed.

The terms "direct-threaded", "indirect-threaded", and "subroutine threaded" are generally invoked to describe the means of breaking the recursion, but they tend to be used differently by different engineers.

Before I get into things, consider a casual assertion I slipped past you in the previous post:

The run-time architecture of a compiled language is a corollary of a virtual machine.

I'm not sure whether this is obvious or not. When I dug into my first fig model, the one for the 6800, it seemed obvious to me. But many people talking about Forth seemed to not even want to talk about the inner-interpreter as a VM. When I started looking at the Pascal virtual machine, I saw all sorts of parallels, and when I started reading the output of the C compilers at school, I saw most of the same essential parallels. (And none of my professors seemed to want to compare run-time architectures with VMs, either.)

  • Stack pointer(s): Both VMs and non-VM run-times usually have one or more stack pointers.
  • Call protocol: Both VMs and non-VM run-times have call protocols.
  • Global variables/parameters: Both VMs and non-VM run-times have them.
  • Local variables/parameters: Likewise.
  • Instruction pointer(s): Again, in both.
  • Scratch registers: In both. 
  • Current procedure/function pointer: Necessary in either if object-like self-inspection is part of the language.
  • Stack frames: of questionable utility in either, more often not present when parameters are separated from return pointers.

Stack frames in C had me scratching my head, at first. I could almost understand why Pascal defined them, because of the existence of local procedures and functions. But C had no such concepts, and still wasted the time to build and tear-down stack frames. The only conclusion I could come to was that the compiler authors wanted the convenience of not having to remember how deeply things were stacked in the current expression evaluation, plus the offset for the local parameters. And it seemed obvious to me that the difficulty in tracking was primarily due to interleaving the stacks.

Should I have considered the possibility that certain well-known system architects simply preferred to have easily crashed stacks?

Anyway, a stack frame is not a required element of a run-time architecture.

Setting aside the call protocols for a bit, let's talk about register mappings.

The fig-Forth model run time includes the following registers:
  • W: pointer to the currently running definition
  • IP: pointer to the next i-code for the VM to execute
  • RP: pointer to the top of the stack of nested calls
  • SP: pointer to the top of the stacked parameters
  • UP: pointer to the base of the user process state table
  • PC: pointer to next machine code to execute (on non-Forth CPUs)
  • scratch registers for intermediate values and operators.

Common C run-times statically map the following registers:

  • IP or PC: pointer to next CPU instruction to perform
  • SP: pointer to the top of the stack of nested calls
  • Link: cached most recent return address, not saved in leaf routines (Not present in all run-times.)
  • Heap pointer: Often implicitly the bottom of global address space
  • Here pointer: Pointer to the state of the currently active object in object-oriented languages
  • Scratch registers for intermediate values and operators

Now we can make some comparisons.

IP is to virtual instructions when a VM is running, and is separate from the processor's PC. But they are the same essential concept.

SP and RP, as mentioned, are tangled concepts for saving the state of a routine or function before it calls another.

W and Link are not the same, even if neither is saved in leaf routines. W is the currently executing definition, Link is the caller.

UP and the heap address/pointer are essentially corollary, even if what they contain is organized differently.

However, ...

Here are some possible register mappings for (fig) Forth vs. certain common C run-times:

On the 68000:

  • A3: W vs. Link (maybe), here pointer, or scratch
  • A4: IP vs. scratch
  • A7: RP vs. interleaved stack pointer
  • A6: Forth SP vs. frame pointer (if present) (Note that MOVEM={PUSH|POP} for all An.)
  • A5:UP vs. heap pointer (if present)
  • A0~A2, D0~D7: scratch

On the 6809:

  • DP[0] (first two bytes in the direct page): W vs. unspecified/scratch
  • Y: IP (save before using for other things) vs. unspecified/scratch
  • S: RP vs. interleaved stack pointer
  • U: Forth SP vs. frame pointer or scratch
  • DP: UP (putting W in the user state table) vs. optional optimization use or per-task variables
  • X, D (A:B) vs. scratch

On the 6800:

  • $F0 in the direct (zero) page: W vs. frame pointer or scratch
  • $F6 in the direct (zero) page: IP vs. scratch
  • SP: RP vs. interleaved stack pointer
  • $F4 in the direct (zero) page: SP vs. scratch
  • $F2: UP vs. heap pointer or scratch
  •  X, A, B: scratch

What I want to point out is that C is really stripped down. This is because C was originally structured to be very lightweight, to use as few of the processor resources as possible, leaving the rest to the compiler or application to make optimal use of. (But current trends in C have been adding things like the here pointer to the run-time model.)

Also, C was designed to, if possible, use no global RAM. This is not possible on the 6800 because there are so few registers, and you have to use scratch RAM to do many things. On the 6801, it becomes possible, if a bit awkward, to move all scratch RAM to an interleaved stack.

I should note that my assignment of the CPU's stack pointer to the return stack pointer and the consequent use of a synthetic stack pointer as the parameter stack pointer is opposite of the fig Forth models. I'm not sure if the engineers who built the 68000 or 6809 versions of the fig models were aware that the MOVEM instruction that does the PUSH and POP on on the 68000 operates equally on all address registers, or that the 6809's U register has it's own PSHU and PULU instructions, with no overhead compared to PSHS and PULS.

With the 6800, it was probable that the the PSH/PUL A/B instructions were desired for the parameter stack, but I have done scratch calculations showing that they provide only minimal advantage. (For me, five percent is minimal when it means something like data and return addresses on the same stack.)

You have to use some sort of global RAM variable for one or the other of the stack pointers on the 6800. If it's in the direct page, the choice balances out a bit.

I'll emphasize this -- I prefer to avoid mixing parameters with return addresses whenever possible.

Now let's look at the call protocols. 

In C, various call protocols have been used. Points where the protocols differ:

  • Saving registers --
    • Calling routine saves the registers it needs preserved, vs.
    • the called routine saves the registers it uses.
  • Building a stack frame or not
  • Caching the most recent caller in a register or not

Likewise, in Forth, various call protocols have been used.

A certain aspect of the call protocol is the primary point of distinction between the three threading types in Forth, but there is still a bit of variation in how the calls are implemented.

I'll initially analyze the 6800 fig model (Find the actual source here, and the assembled listing here.), then extrapolate from there. In this analysis, I'll use the fig Forth register assignments, with all VM registers in the direct (zero) page except for SP:

  • W    RMB    2    the instruction register points to 6800 code
  • IP    RMB    2    the instruction pointer points to pointer to 6800 code
  • RP    RMB    2    the return stack pointer
  • UP    RMB    2    the pointer to base of current user's 'USER' table  ( altered during multi-tasking ) 
  • N    RMB    10    used as scratch by various routines
  • * SP is S

(For an in-depth look at how this plays out using a synthetic stack for the parameters and the CPU stack for RP, see here.)

The inner interpreter of the virtual machine looks like this:

NEXT
    LDX    IP
    INX        pre-increment fetch from i-code list in definition
    INX
    STX    IP   Update it.
NEXT2
    LDX    0,X    Get the i-code which points to a pointer to CPU code
NEXT3
    STX    W      This is the same as the label in the symbol table.
    LDX    0,X    Get pointer to executable code.
NEXTGO
    JMP    0,X    Jump to executable code.

The definition header consists of a symbol table entry like this:

    HEADER
LABEL CODE_POINTER
    DEFINITION
The header consists of the name that the Forth command line interpreter will recognize, massaged a bit, and a link to the previous definition. (I'll explain some of that below.)

The CODE_POINTER is a pointer to machine language code the CPU can execute directly.

For a low-level definition, the definition consists of the machine code that the CPU can execute, and this is where the CODE_POINTER points. The code ends with a jump back to NEXT (or some other appropriate place that leads back to NEXT).

For a "high-level" (non-leaf) definition, the CODE_POINTER points to a nesting routine, and the definition consists of the list of definition addresses -- as i-codes for the virtual machine inner interpreter as described in part one of this rant. This list will be terminated by an i-code that unnests the VM and returns to the caller.

For an easy example of a low-level (leaf) definition, we can look at the code to add a double integer, starting at line 1008 in the assembled listing. 

Note that the header starts with the length of the symbol name for Forth masked with some mode bits, followed by all but the last of the characters of the name, then the last character with its high bit set. ("$" says hexadecimal, ASCII for "+" is hexadecimal 4B.) The mode/length byte has its high bit set, as well, to brace the name. The link to the previous definition in the table, PLUS, is adjusted to point to that definition's name: two bytes of link, one byte of length, and one byte of name makes 4. 

"*+" means "address of here plus 2", which address is where the machine code starts:


  13FD 82          	FCB	$82
  13FE 44          	FCC	1,D+
  13FF AB          	FCB	$AB
  1400 13 ED       	FDB	PLUS-4
  1402 14 04       DPLUS	FDB	*+2
  1404 30          	TSX
  1405 0C          	CLC
  1406 C6 04       	LDA B	#4
  1408 A6 03       DPLUS2	LDA A	3,X
  140A A9 07       	ADC A	7,X
  140C A7 07       	STA A	7,X
  140E 09          	DEX
  140F 5A          	DEC B
  1410 26 F6       	BNE	DPLUS2
  1412 31          	INS
  1413 31          	INS
  1414 31          	INS
  1415 31          	INS
  1416 7E 10 34    	JMP	NEXT

For an easy example of a definition that consists of i-codes, we can look at a convenience definition that adds 2 to the top integer on stack, starting at line 1535. The listing is a little awkward -- it's hard to see that TWOP labels the address 171D, and that the six bytes following are the label values of DOCOL (1525), TWO (15B5), and PLUS (13F1). (Remember, the 6800 is most significant byte first, so you don't have to reverse the bytes in your mind.)

So the CODE_POINTER for "2+" is DOCOL, and the i-code list ends with SEMIS.


  1718 82          	FCB	$82
  1719 32          	FCC	1,2+
  171A AB          	FCB	$AB
  171B 17 0B       	FDB	ONEP-5
  171D 15 25 15 B2 13 F1 
                        TWOP	FDB	DOCOL,TWO,PLUS
  1723 13 67       	FDB	SEMIS

Let's look at DOCOL, starting at line 1206. It's part of the mixed definition COLON, and has no header of its own:


  1525 DE F4       DOCOL	LDX	RP	make room in the stack
  1527 09          	DEX
  1528 09          	DEX
  1529 DF F4       	STX	RP
  152B 96 F2       	LDA A	IP
  152D D6 F3       	LDA B	IP+1	
  152F A7 02       	STA A	2,X	Store address of the high level word
  1531 E7 03       	STA B	3,X	that we are starting to execute
  1533 DE F0       	LDX	W	Get first sub-word of that definition
  1535 7E 10 36    	JMP	NEXT+2	and execute it

We can see how it decrements RP to save the current IP, saves it, loads the address that NEXT just stored in W, then jumps into the right place in the NEXT inner interpreter to save it as the new IP and store the appropriate new W.

To brace our understanding of this, let's look at SEMIS, starting at line 896:


   1362 82          	FCB	$82
   1363 3B          	FCC	1,;S
   1364 D3          	FCB	$D3
   1365 13 52       	FDB	RPSTOR-6
   1367 13 69       SEMIS	FDB	*+2
   1369 DE F4       	LDX	RP
   136B 08          	INX
   136C 08          	INX
   136D DF F4       	STX	RP
   136F EE 00       	LDX	0,X	get address we have just finished.
   1371 7E 10 36    	JMP	NEXT+2	increment the return address & do next word

We can see that it is a low-level leaf. It increments RP, gets the saved IP, and jumps into the right place in NEXT to store it back in IP and continue where the caller left off.

This threading of functionality through low-level and high-level definitions is what the Forth community calls "threaded".

And the above model is an example of indirect threading, where the i-codes are pointers to pointers to code.

In case you're wondering, no, this is not the only way to do indirect threading. It's just one example.

Now, what about direct threading?

We'll stick with the same register model. And we'll start with the routine for double add, since it's sort-of familiar. The header structure will be the same, up to the label.

DPLUS TSX ; index the parameters
CLC ; so we can do this in a loop
LDA B #4 ; bytes to add, start with least significant
DPLUSL LDA A 3,X ; right-hand term
ADC A 7,X ; left-hand term
STA A 7,X ; overwrite left-hand term
DEX ;
DEC B ; done?
BNE DPLUSL ; back for more
INS ; Deallocate right-hand term.
INS
INS
INS
JMP NEXT 

the only thing that has changed is that the CODE_POINTER is missing in SEMIS. So how do the i-code lists get their starts? Let's look at 2+:

TWOP JMP DOCOL
FDB TWO,PLUS,SEMIS

Now the i-code lists start with a little machine language code. On the 6801, DOCOL could be located in the direct page, and the jump would only be two bytes, FWTW. But you'd still be starting an i-code list with something that doesn't even look like an i-code. That adds a bump in designing debuggers, and in code analysis of the i-code lists.

Let's see how NEXT changes to support this:

NEXT LDX IP
INX ; We can still do pre-increment mode
INX
STX IP
NEXT2 LDX 0,X ; get W which points to definition to be done
NEXT3 STX W
JMP 0,X ; Just jump right to it.

This looks like it ought to be faster, by a small amount. What happens to DOCOL and SEMIS, then? Interestingly enough, they don't have to change, except for losing the CODE_POINTER. 

So why use indirect threading?

Basically, it keeps the i-code list pure, which simplifies debugging tools and such. Also, indirect threading helps avoid some of the problems your developers' tools may make for you, when they try to help you. (Yet another topic.)

Again, this is not the only way to do direct-threading.

How about subroutine threading?

In subroutine threading, the definitions are called by subroutine. This can be done both indirect-threaded and direct-threaded, but, typically, it is done direct-threaded, in the idea that it is faster. NEXT looks like this:

NEXT LDX IP
INX ; We can still do pre-increment mode
INX
STX IP
NEXT2 LDX 0,X ; get W which points to definition to be done
NEXT3 STX W
* LDX 0,X ; if doing indirect
JSR 0,X ; Use a native call.

To match this, low-level definitions end in a RTS instead of a JMP NEXT. The call and return actually take a bit more time than JMP, having to save and restore data on the CPU stack. 

DOCOL doesn't have to change, and SEMIS ends in a RTS. Certain definitions we haven't looked at change because of return addresses that are now on the parameter stack. (Using the CPU's call stack for the parameter stack causes the greater ripple effect here.)

Again, these are not the only ways of doing subroutine threading. In particular, since the call in the direct-threaded model actually saves W on the return stack, explicitly having a W register in the VM and storing X to it at NEXT3 can be done away with. Such an approach does requires changes to DOCOL, however.

Some daredevils pervert the CPU's stack pointer into the IP and NEXT becomes a CPU return. I say daredevil, because, if the CPU gets any sort of interrupt, the interrupt walks all over the code. (Most CPUs save interrupt state on the call stack.) If you're that desperate to use an IP register with auto-increment mode, use a CPU that has a valid auto-increment mode for a non-stack register.

Just for curiosity's sake, let's look at non-subroutine indirect threading on the 6809, going with my recommendation of using the U stack for parameters this time. The header will not change. Registers in the virtual machine will be assigned as above:

  • W: X (Save before using for other things, if you need W.)
  • IP: Y (save before using for other things and restore before NEXT)
  • RP: S
  • SP: U
  • UP: DP

The examples, as with the 6800:

DPLUS   FDB *-2
    LDD   6,U  ; left-hand
    ADDD   2,U  ; right-hand
    STD   6,U
    LDD   4,U  ; left-hand
    ADCB   1,U ; No ADCD, do it by bytes.
    ADCA   ,U
    LEAU   4,U ; deallocate before store to save a cycle.
    STD   ,U
    JMP   NEXT
*    JMP   [,Y++] ; This could be NEXT on the 6809, if ignoring W
*

 DOCOL, NEXT, and SEMIS are significantly optimized:

DOCOL    LDX   ,Y++   ; using post-increment instead of pre-
    PSHS   Y    ; save pointer to next
    LDY   ,Y    ; Get new IP
NEXT   LDX   ,Y++  ; using post-increment instead of pre-
    JMP   [,X]
*
SEMIS    PULS   Y
    BRA NEXT

How about the 68000? It gives us a 32-bit add naturally, so we'll define a 64-bit double add. Using the registers as I suggest above,

DPLUS   DC.L   *+4
    MOVEM   (A6),D0/D1/D2/D3
    ADD.L   D1,D3   ; less significant 32 bits
    ADDX.L   D0,D2   ; more significant 32 bits
    LEA   8(A6),A6   ; deallocate right-hand term
    JMP   NEXT
And

DOCOL    MOVE.L   (A4)+,A3   ; using post-increment instead of pre-
    MOVE.L    A4,-(A7)   ; save pointer to next
    MOVE.L    A3,A4    ; Get new IP
NEXT   (A4)+,A3  ; using post-increment instead of pre-
    MOVE.L    (A3),A2
    JMP   (A2)
*
SEMIS    MOVE.L (A7)+,A4
    BRA NEXT
The shift from indirect-threaded to direct, and the use of subroutine-threading would follow pretty much as the shift in the 6800, modulus the advantage of having the registers in actual registers instead of in memory.

There is one more step I mentioned in passing in part one of this rant. If using native calls, a certain degree of speed optimization can be obtained by flattening an i-code list. Instead of starting the definition to be optimized this way with a jump to the inner interpreter, the definition can be compiled as a series of calls.

On the 6809, 2+ can be changed from

TWOP    JMP DOCOL
        FDB TWO,PLUS,SEMIS 

to 

TWOP    JSR     TWO
    JSR     PLUS
    RTS
And from there, optimizing to

TWOP    LDD    #2
    ADDD    ,U
    STD    ,U
    RTS

becomes almost trivial.

(I also discuss optimizations on the 6800 a little in the synthetic stack examples here: https://defining-computers.blogspot.com/2020/10/6800-example-vm-with-synthetic.html.)

Caveat: It's way past my bedtime again, and I may have glaring mistakes in the above. If so, leave me comments, please.

Why is this important?

Using the fig Forth VM, it is a bit difficult to share the Forth words as library code for compiled languages such as C. Using the native CPU calls instead, it becomes possible to share, if the compiler for the other language uses a split stack.

Thursday, October 22, 2020

Computer Languages -- Interpreted vs. Compiled Forth?

[JMR202010251132:

Comments on a post by Peter Forth in the Forth2020 Users-Group on FB set me off on a long response that turned into this webrant on the distinctions between

Interpreter vs. compiler? -- Interpreted vs. Compiled? 

Forth?

][JMR202010251132 -- edited for clarification.]

This is a particularly tangled intersection of concepts and jargon.

Many early BASICs directly interpreted the source code text. If you typed

PRINT "HELLO "; 4*ATN(1)

and hit ENTER, the computer would parse (scan) along the line of text you had typed and find the space after the "PRINT", and go to the symbol table (the BASIC interpreter's dictionary) and look it up. There it would find some routines that would prepare to print something out on the screen (or other active output device).

Then it would keep parsing and find the quoted 

"HELLO " 

followed by the semicolon, and it would prepare the string "HELLO " for printing, probably by putting the string into a print buffer.

Continuing to scan, it would find the numeric text "4", followed by the asterisk. It would recognize it as a number and convert the numeric text to the number 4, then, since asterisk is the multiplication symbol in BASIC, put both the number and the multiplication operation somewhere to remember them. 

Then it would scan the symbol "ATN" followed by the opening parenthesis character. Looking this up in the symbol table, it would find a routine to compute the arctangent of a number, and remember that waiting function call, along with the fact that it was looking for the closing parenthesis character.

Then it would scan the text "1", followed by the closing parenthesis. Recognizing the text as a number, it would convert it to the number 1. Then it would act on the closing parenthesis and call the saved function, passing it the 1 to work on. The function would return the result,  0.7853982, and save it away after the 4 stored earlier.

Then it would see that the line had come to an end, and it would go back and see what was left to do. It would find the multiplication, and multiply 4 x 0.7853982, giving 3.1415927 (which needs a little explaining), and remember the result. Then it would see that it was in the middle of PRINTing, and convert the number to text, put it in the buffer following the string, and print the buffer to the screen, something like

HELLO  3.1415927

I know that seems like a lot of work to go to. Maybe it would not help you to know I'm only giving you the easy overview. There are lots more steps I'm not mentioning.

(So, you're wondering why 4 x 0.7853982 is 3.1415927, not 3.1415928? Cheap calculators are inexact after about four to eight digits of decimal fraction. BASIC languages tended to implement a cheap calculator -- especially early BASIC languages. But, just for the record, written with a few bits more accuracy, π/4 ≒ 0.78539816; π ≒ 3.14159266. The computer often keeps more accuracy than it shows.)

Interpreted vs. compiled actually indicates a spectrum, not a binary classification. What I have described above is one extreme end of the spectrum, referred to in such terms as "pure source text interpreter".

Other approaches to the BASIC language were taken. Some were pure compiled languages. Instead of acting immediately on each symbol as it parses, compilers save the corresponding code in  the CPU's native machine language away in a file, and the user has to call the compiled program back up to run it -- after the compiler has finished checking that the code follows the grammar rules and has finished saving all the generated machine code. (Again, there are steps I am not mentioning at this point.)

In the case of the above line of BASIC, the resultant machine code might look something like the following, if the compiler compiles to object code for the 6809 CPU: 

48454C4C4F20
00
40800000
3F800000
308CEE
3610
17CFE9
308CED
3610
308CEC
3610
17D828
17D801
17CFEE
17CFF7

That isn't very easily read by humans. Here is what the above looks like in 6809 assembly language:

1000                  00007         S0001
1000 48454C4C4F20     00008             FCC 'HELLO '
1006 00               00009             FCB 0
1007                  00010         FP0001
1007 40800000         00011             FCB $40,$80,00,00 ; 4
100B                  00012         FP0002
100B 3F800000         00013             FCB $3F,$80,00,00 ; 1
100F                  00014         L0001
100F 308CEE           00015             LEAX S0001,PCR
1012 3610             00016             PSHU X
1014 17CFE9           00017             LBSR BUFPUTSTR
1017 308CED           00018             LEAX FP0001,PCR
101A 3610             00019             PSHU X
101C 308CEC           00020             LEAX FP0002,PCR
101F 3610             00021             PSHU X
1021 17D828           00022             LBSR ATAN
1024 17D801           00023             LBSR FPMUL
1027 17CFEE           00024             LBSR BUFPUTFP
102A 17CFF7           00025             LBSR PRTBUF

Think you could almost read that? You see the word "HELLO " in there, and you can guess that 

  1. FP0001 and FP0002 are the floating point encodings for 4 and 1, respectively. 
  2. LEAX calculates a pointer to the value given to it, and 
  3. LBSR calls subroutines which
    1. put the string in the print buffer, 
    2. calculate the arctangent, 
    3. multiply the arctangent (by the saved 4.0), 
    4. convert the result to text and put it in the print buffer, 
    5. and send the print buffer to the screen (or other current output device).

Since the CPU interprets the machine language directly, the compiled code is going to run very fast. 

With something this short, we don't really care about how fast it is, but with a really long program we might care a lot.

On the other hand, with something as short as this, we are less interested in how fast it runs than in the effort and time it takes to prepare and run the code. 

Compiled languages add a compile step between typing the code in and running it. If we want to change something, we have to edit the source code and run it through the compiler again. 

On modern computers with certain integrated developer environments, that's actually less trouble than it sounds, but without those IDEs, or on slower computers, the pure text interpreter is much easier to use for short programs.

On eight bit computers, those fancy IDEs took too much of the computer's resources, because the machine code of a CPU can be very complicated.

Hmm. The example of compiling to the 6809 is a bit counter to my purpose here, because the architecture and machine code of the 6809 were at once quite simple and well-matched to the programs we write. 

Let's see what that line of code would look like compiled on a "modern" CPU. I'm going to cheat a little and show you what it looks like written in C, and what it looks like compiled from C, because I don't want to take the time to install either FreeBasic or Gambas in my workstation and figure out how to get a look at the compiled output. Here's the C version, wrapped in the mandatory main() procedure:

/* Using the C compiler as a calculator.
*/

#include <stdlib.h>
#include <math.h>
#include <stdio.h>

int main( int argc, char *argv[] )
{
  printf( "HELLO %g\n", 4 * atan( 1 ) );
}

That's a lot of wrapper just to do calculations, don't you think?

(Not all compiled languages require the wrapping main function to be explicitly declared, but many do.)

Well, here's the assembly language output on my Ubuntu (Intel64 perversion of AMD64 evolution of 'x86 CPU) workstation: 

    .file    "printsample.c"
    .text
    .section    .rodata
.LC1:
    .string    "HELLO %g\n"
    .text
    .globl    main
    .type    main, @function
main:
.LFB5:
    .cfi_startproc
    pushq    %rbp
    .cfi_def_cfa_offset 16
    .cfi_offset 6, -16
    movq    %rsp, %rbp
    .cfi_def_cfa_register 6
    subq    $32, %rsp
    movl    %edi, -4(%rbp)
    movq    %rsi, -16(%rbp)
    movq    .LC0(%rip), %rax
    movq    %rax, -24(%rbp)
    movsd    -24(%rbp), %xmm0
    leaq    .LC1(%rip), %rdi
    movl    $1, %eax
    call    printf@PLT
    movl    $0, %eax
    leave
    .cfi_def_cfa 7, 8
    ret
    .cfi_endproc
.LFE5:
    .size    main, .-main
    .section    .rodata
    .align 8
.LC0:
    .long    1413754136
    .long    1074340347
    .ident    "GCC: (Ubuntu 7.5.0-3ubuntu1~18.04) 7.5.0"
    .section    .note.GNU-stack,"",@progbits

Now, with a bit of effort, I can read that. I can see where it's setting up a stack frame on entry and discarding it on exit. I can see where the compiler has pre-calculated the constant, so that all that is left to do at run-time is load the result and print it. But, even though you can see the string "HELLO " and the final call to printf(), if you can decipher the stuff in between without explanation, you probably already know anything I can tell you about this stuff. 

The IDE running on your computer has to keep track of stuff like this very quickly, in order to make it easy for you to change something and get quick results. (Back in the 1980s, the Think C compiler for the 68000 (Apple Macintosh) was able to do stuff like this, in no small part because the 68000, like it's little brother the 6809, is pretty powerful even though it's relatively simple.)

So, back in the 1970s and 1980s, companies that weren't, for various reasons, willing to use either the 6809 or 68000 in their designs, wanted something a little simpler than the CPU to compile to, so that the IDE (and the humans) could keep track of what was happening.

So they devised "virtual machines" that interpreted an intermediate code between the complexity of the CPU's raw machine code and the readability of pure text. These virtual machines also made it easier to have compilers that were portable between various processors, but that's another topic.

The intermediate codes these VMs ran were known as "i-codes", but were often called "p-codes" (for Pascal i-codes) or "byte-codes" (because many of them, especially Java, were byte-oriented encodings). (And, BTW, "P-code" was also an expression used at the time to indicate pseudo-coding, which can be a bit confusing, but that's another topic.)

If you want an idea what an i-code looked like, one possible i-code using prefix 3-tuples and converted to human-readable form might look like

string 'HELLO '
putstr stdout s0001
float 4
float 1
library    arctan fp0002
multiply fp0001 res0001
putfloat stdout res0002
printbuffer stdout

A stacked mixed-tuple postfix i-code would also be possible:

'HELLO ' putstr
4.0
1.0
arctan
multiply
putfloat
stdout printbuffer

There were versions of BASIC that ran on VMs. The Basic compilers compiled the source code to the i-codes, and the VM interpreted the i-codes. 

Note that we have added a level of interpretation between the source text and the CPU.

You might think that the 6809 wouldn't need an i-code interpreted BASIC, but a byte-code could be designed to be more compact even than 6809 machine code. 

Basic09 was basically developed at the same time as the processor itself was designed. The 6809 was designed to run the  Basic09 VM very efficiently, and it did. But it was a compiled language, in the sense that there was no command-line interpreter built-in. (You could build or buy one, but it wasn't built-in.) It was source-code compiled, i-code interpreted. (Analyzing the functionality of Basic09 is also yet another topic, but it worked out well, partly because of the host OS-9 operating system.)

Even now, there are a number of languages that (mostly) run in a VM, for various reasons. The modular VM of Java, with the division between the CPU and the language, can make the language safer to use, which is yet another topic. (Except that now Java has had so much added to it, that .... Which is another topic, indeed.)

Which brings me to Forth.

Early Forth languages tended to be i-code interpreted VMs, for simplicity and portability. The i-code was very efficient. Not byte-code, at least not originally, but i-code. 

(Specifically, Forth built-in operations were essentially functions, and, when defining a new function, the function definition would be compiled as a list of addresses of the functions to call. A newly-defined function would also be callable by its address in the same way as built-in functions. How that is done involves something the Forth community calls threaded code, which is, ahem, another topic, partially addressed here: https://defining-computers.blogspot.com/2020/10/forth-threading-is-not-process-threads.html.)

Forth interpreters include a command-line interpretation mode, which is made easier because of the i-code interpretation. 

(In a sense, the interpreter mode is much like a directly interpreted BASIC interpreter with a giant case statement, where new definitions implicitly become part of the giant case statement.)

The built-in i-codes include commands that can be invoked as post-fix calculator commands, directly from the command-line.

The above example, entered at the command line of gforth, which is a Forth that knows floating-point, would look like this:

4.0e0   1.0e0   fatan   f*   f.

(giving the result 3.14159265358979.)

Uh, yes, I know, in terms of readability, that only looks marginally improved over 6809 assembler source. Maybe.

Lemme 'splain s'more!

That's because Forth is, essentially, the i-code of it's VM. Not byte-code, but address-word-code VM.

Forth has a postfix grammar, something like the RPN of Hewlett Packard calculators. So the first two text blobs are 4.0 x 100 and 1.0 x 100 in a modified scientific notation, pushed on the floating point stack in that order. Fatan is the floating point arctangent function, which operates on the topmost item on the floating point stack, 1.0. It returns the result -- 0.785398163397448, or a quarter of π -- in the place of the 1.0e0. 

'F*' is not a swear word, it's the floating point multiply, and it multiplies the top two items on the floating point stack, resulting in (an approximation of) π. 'F.'  prints the top floating point number to the current output device, in this case, the screen.

The stack, in combination with the post-fix notation, allows most of the symbols the user types in from the command-line to have a one-to-one correspondence with the i-codes themselves.

The 6809 and 68000 have an interesting quality in common (almost, but not quite, shared by the 8086) -- Both the 6809 and the 68000 have an instruction set that includes op-codes that can be mapped one-to-one or one-to-two with a significant subset of the Forth primitives.

There are have been other processors designed specifically to be mapped one-to-one with all the important Forth leaf-level primitives. The Forth Interest Group lists a few of the current ones on one of their web pages.

Thus, Forth is a bit unusual, even among byte-code and i-code languages. It partakes of both ends of the spectrum, and of the middle, all at once, being both interpreted at the text level and compiled at the machine-code level, as well as operating at the intermediate level (although not at all levels at once in every implementation).

Now, before we think we have covered the bases, we should note that all modern compilers compile to an intermediate form, as a means of separating the handling of the source language from the handling of the target CPU.

Forth might actually define one of the possible intermediate forms a compiler might compile to, if the compiler compiles to a split stack run-time model instead of a linked-list-of-interleaved-frames-in-a-single-stack-model. 

(Note here that the run-time model of compiled languages is essentially a close corollary of a virtual machine.)

An interleaved single-stack model mixes local variables, parameters, and return addresses on the same stack. This means that a called function can walk all over both the link to the calling frame and the address in the calling code to return to, if the programmer is only a little bit careless. When that happens, the program usually loses complete track of what it's doing, and may begin doing many kinds of random and potentially damaging things.

This is where most of the stack vulnerabilities and stack attack vectors that you hear about in software come from.

Some engineers consider the single stack easier to manage and maintain than a split stack. You only have to ask the operating system to allocate one stack region, and you only have to worry about overflow and underflow in that region.

But interleaving the return addresses with the parameters and local variables comes at a trade-off cost of having to worry about frame bounds in every function, in addition to the cost of maintaining the frame.

With a (correctly) split stack, there is no need to maintain frames in most run-times because they maintain themselves, and it takes more than carelessness to overwrite the return address (and optional frame pointer for those languages which need one).

Given this, one might expect that most modern run-times would prefer the problems of maintaining a split stack over the problems of maintaining a single stack. 

But memory was tight in the 1980s (or so we thought), and the single stack was the model the industry grew up with. (Maybe you can tell, but I believe this was a serious error.)

(By the way, this is one of the key places where the 8086 misses being able to implement a near one-to-one mapping -- segmentation biases it towards the single-stack model. BP shares a segment register with SP. If BP had its own segment register, you could put the return address stack in it's own somewhat well-protected segment, but it doesn't, so you can't.)

Especially with CPUs that map well to the Forth model, Forth interpreters can be designed to compile to machine code, and run without an intermediary virtual machine. [JMR202010222211: I'm actually working -- a little at a time -- on a project that will do that, here: https://osdn.net/projects/splitstack-runtimelib/.]

One more topic that needs to be explicitly addressed: Most languages have only limited ability to extend their symbol tables, especially at run-time. Most text-only interpreters have severe limits to extending their symbol tables -- in some cases just ten possible user functions named fn0 to fn9, in some cases, no extensions at all.

I've implied this above, but, by design, Forth interpreters do not have such limits. Symbols defined by the programmer have pretty much equal standing with pre-defined symbols. This makes Forth even less like most interpreted languages.

Yes, Forth is often an interpreted language, but not like other interpreted languages.

And it's after midnight here, and I keep drifting off, so this rant ends here.

[JMR202010251351:

Discussing of threading added in a separate post: https://defining-computers.blogspot.com/2020/10/forth-threading-is-not-process-threads.html, also linked above.

]