In comp.arch.fpga Arlet Ottens <usenet+5@c-scape.nl> wrote:> On 04/04/2013 01:16 PM, Albert van der Horst wrote:>>>> You mentioned code density. AISI, code density is purely a CISC >>>> concept. They go together and are effectively inseparable.>>> They do go together, but I am not so sure that they are inseperable.>>> CISC began when much coding was done in pure assembler, and anything >>> that made that easier was useful. (One should figure out the relative >>> costs, but at least it was in the right direction.)>> But, of course, this is a fallacy. The same goal is accomplished by >> macro's, and better. Code densitity is the only valid reason.> Speed is another valid reason.Presumably some combination of ease of coding, speed, and also Brooks' "Second System Effect". Paraphrasing from "Mythical Man Month" since I haven't read it recently, the ideas that designers couldn't implement in their first system that they designed, for cost/efficiency/whatever reasons, come out in the second system. Brooks wrote that more for OS/360 (software) than for S/360 (hardware), but it might still have some effect on the hardware, and maybe also for VAX. There are a number of VAX instructions that seem like a good idea, but as I understand it ended up slower than if done without the special instructions. As examples, both the VAX POLY and INDEX instruction. When VAX was new, compiled languages (Fortran for example) pretty much never did array bounds testing. It was just too slow. So VAX supplied INDEX, which in one instruction did the multiply/add needed for a subscript calcualtion (you do one INDEX for each subscript) and also checked that the subscript was in range. Nice idea, but it seems that even with INDEX it was still too slow. Then POLY evaluates a whole polynomial, such as is used to approximate many mathematical functions, but again, as I understand it, too slow. Both the PDP-10 and S/360 have the option for an index register on many instructions, where when register 0 is selected no indexing is done. VAX instead has indexed as a separate address mode selected by the address mode byte. Is that the most efficient use for those bits? -- glen
MISC - Stack Based vs. Register Based
Started by ●March 29, 2013
Reply by ●April 4, 20132013-04-04
Reply by ●April 4, 20132013-04-04
In comp.arch.fpga Syd Rumpo <usenet@nononono.co.uk> wrote: (snip)> Can you achieve as fast interrupt response times on a register-based > machine as a stack machine? OK, shadow registers buy you one fast > interrupt, but that's sort of a one-level 2D stack.If you disable interrupts so that another one doesn't come along before you can save enough state for the first one, yes. S/360 does it with no stack. You have to have some place in the low (first 4K) address range to save at least one register. The hardware saves the old PSW at a fixed (for each type of interrupt) address, which you also have to move somewhere else before enabling more interrupts of the same type.> Even the venerable RTX2000 had an impressive (IIRC) 200ns interrupt > response time.-- glen
Reply by ●April 4, 20132013-04-04
On 04/04/2013 02:49 PM, glen herrmannsfeldt wrote:> In comp.arch.fpga Syd Rumpo <usenet@nononono.co.uk> wrote: > > (snip) >> Can you achieve as fast interrupt response times on a register-based >> machine as a stack machine? OK, shadow registers buy you one fast >> interrupt, but that's sort of a one-level 2D stack. > > If you disable interrupts so that another one doesn't come along > before you can save enough state for the first one, yes. > > S/360 does it with no stack. You have to have some place in the > low (first 4K) address range to save at least one register. > The hardware saves the old PSW at a fixed (for each type of > interrupt) address, which you also have to move somewhere else > before enabling more interrupts of the same type.ARM7 is similar. PC and PSW are copied to registers, and further interrupts are disabled. The hardware does not touch the stack. If you want to make nested interrupts, the programmer is responsible for saving these registers. ARM Cortex has changed that, and it saves registers on the stack. This allows interrupt handlers to be written as regular higher language functions, and also allows easy nested interrupts. When dealing with back-to-back interrupts, the Cortex takes a shortcut, and does not pop/push the registers, but just leaves them on the stack.
Reply by ●April 4, 20132013-04-04
On 4/3/13 11:31 PM, Arlet Ottens wrote:> On 04/04/2013 10:38 AM, Syd Rumpo wrote: >> On 29/03/2013 21:00, rickman wrote: >>> I have been working with stack based MISC designs in FPGAs for some >>> years. All along I have been comparing my work to the work of others. >>> These others were the conventional RISC type processors supplied by the >>> FPGA vendors as well as the many processor designs done by individuals >>> or groups as open source. >> >> <snip> >> >> Can you achieve as fast interrupt response times on a register-based >> machine as a stack machine? OK, shadow registers buy you one fast >> interrupt, but that's sort of a one-level 2D stack. >> >> Even the venerable RTX2000 had an impressive (IIRC) 200ns interrupt >> response time. > > It depends on the implementation. > > The easiest thing would be to not save anything at all before jumping to > the interrupt handler. This would make the interrupt response really > fast, but you'd have to save the registers manually before using them. > It would benefit systems that don't need many (or any) registers in the > interrupt handler. And even saving 4 registers at 100 MHz only takes an > additional 40 ns.The best interrupt implementation just jumps to the handler code. The implementation knows what registers it has to save and restore, which may be only one or two. Saving and restoring large register files takes cycles!> If you have parallel access to the stack/program memory, you could like > the Cortex, and save a few (e.g. 4) registers on the stack, while you > fetch the interrupt vector, and refill the execution pipeline at the > same time. This adds a considerable bit of complexity, though. > > If you keep the register file in a large memory, like a internal block > RAM, you can easily implement multiple sets of shadow registers. > > Of course, an FPGA comes with flexible hardware such as large FIFOs, so > you can generally avoid the need for super fast interrupt response. In > fact, you may not even need interrupts at all.Interrupts are good. I don't know why people worry about them so! Cheers, Elizabeth -- ================================================== Elizabeth D. Rather (US & Canada) 800-55-FORTH FORTH Inc. +1 310.999.6784 5959 West Century Blvd. Suite 700 Los Angeles, CA 90045 http://www.forth.com "Forth-based products and Services for real-time applications since 1973." ==================================================
Reply by ●April 4, 20132013-04-04
On 4/4/2013 8:44 AM, glen herrmannsfeldt wrote:> In comp.arch.fpga Arlet Ottens<usenet+5@c-scape.nl> wrote: >> On 04/04/2013 01:16 PM, Albert van der Horst wrote: > >>>>> You mentioned code density. AISI, code density is purely a CISC >>>>> concept. They go together and are effectively inseparable. > >>>> They do go together, but I am not so sure that they are inseperable. > >>>> CISC began when much coding was done in pure assembler, and anything >>>> that made that easier was useful. (One should figure out the relative >>>> costs, but at least it was in the right direction.) > >>> But, of course, this is a fallacy. The same goal is accomplished by >>> macro's, and better. Code densitity is the only valid reason.Albert, do you have a reference about this?>> Speed is another valid reason. > > Presumably some combination of ease of coding, speed, and also Brooks' > "Second System Effect". > > Paraphrasing from "Mythical Man Month" since I haven't read it recently, > the ideas that designers couldn't implement in their first system that > they designed, for cost/efficiency/whatever reasons, come out in the > second system. > > Brooks wrote that more for OS/360 (software) than for S/360 (hardware), > but it might still have some effect on the hardware, and maybe also > for VAX. > > There are a number of VAX instructions that seem like a good idea, but > as I understand it ended up slower than if done without the special > instructions. > > As examples, both the VAX POLY and INDEX instruction. When VAX was new, > compiled languages (Fortran for example) pretty much never did array > bounds testing. It was just too slow. So VAX supplied INDEX, which in > one instruction did the multiply/add needed for a subscript calcualtion > (you do one INDEX for each subscript) and also checked that the > subscript was in range. Nice idea, but it seems that even with INDEX > it was still too slow. > > Then POLY evaluates a whole polynomial, such as is used to approximate > many mathematical functions, but again, as I understand it, too slow. > > Both the PDP-10 and S/360 have the option for an index register on > many instructions, where when register 0 is selected no indexing is > done. VAX instead has indexed as a separate address mode selected by > the address mode byte. Is that the most efficient use for those bits?I think you have just described the CISC instruction development concept. Build a new machine, add some new instructions. No big rational, no "CISC" concept, just "let's make it better, why not add some instructions?" I believe if you check you will find the term CISC was not even coined until after RISC was invented. So CISC really just means, "what we used to do". -- Rick
Reply by ●April 4, 20132013-04-04
In comp.arch.fpga rickman <gnuarm@gmail.com> wrote: (snip, then I wrote)>> Then POLY evaluates a whole polynomial, such as is used to approximate >> many mathematical functions, but again, as I understand it, too slow.>> Both the PDP-10 and S/360 have the option for an index register on >> many instructions, where when register 0 is selected no indexing is >> done. VAX instead has indexed as a separate address mode selected by >> the address mode byte. Is that the most efficient use for those bits?> I think you have just described the CISC instruction development > concept. Build a new machine, add some new instructions. No big > rational, no "CISC" concept, just "let's make it better, why not add > some instructions?"Yes, but remember that there is competition and each has to have some reason why someone should by their product. Adding new instructions was one way to do that.> I believe if you check you will find the term CISC was not even coined > until after RISC was invented. So CISC really just means, "what we used > to do".Well, yes, but why did "we used to do that"? For S/360, a lot of software was still written in pure assembler, for one reason to make it faster, and for another to make it smaller. And people were just starting to learn that people (writing software) are more expensive that machines (hardware). Well, that is about the point that it was true. For earlier machines you were lucky to get one compiler and enough system to run it. And VAX was enough later and even more CISCy. -- glen
Reply by ●April 4, 20132013-04-04
On Apr 4, 10:04=A0pm, rickman <gnu...@gmail.com> wrote:> On 4/4/2013 8:44 AM, glen herrmannsfeldt wrote: > > > In comp.arch.fpga Arlet Ottens<usene...@c-scape.nl> =A0wrote: > >> On 04/04/2013 01:16 PM, Albert van der Horst wrote: > > >>>>> You mentioned code density. =A0AISI, code density is purely a CISC > >>>>> concept. =A0They go together and are effectively inseparable. > > >>>> They do go together, but I am not so sure that they are inseperable. > > >>>> CISC began when much coding was done in pure assembler, and anything > >>>> that made that easier was useful. (One should figure out the relativ=e> >>>> costs, but at least it was in the right direction.) > > >>> But, of course, this is a fallacy. The same goal is accomplished by > >>> macro's, and better. Code densitity is the only valid reason. > > Albert, do you have a reference about this? > >Let's take two commonly used S/360 opcodes as an example of CISC; some move operations. MVC (move 0 to 255 bytes) MVCL (move 0 to 16M bytes). MVC does no padding or truncation. MVCL can pad and truncate, but unlike MVC will do nothing and report overflow if the operands overlap. MVC appears to other processors as a single indivisible operation; every processor (including IO processors) sees storage as either before the MVC or after it; it's not interruptible. MVCL is interruptible, and partial products can be observed by other processors. MVCL requires 4 registers and their contents are updated after completion of the operation; MVC requires 1 for variable length moves, 0 for fixed and its contents are preserved. MVCL has a high code setup cost; MVC has none. Writing a macro to do multiple MVCs and mimic the behaviour of MVCL? Why not? It's possible, if a little tricky. And by all accounts, MVC in a loop is faster than MVCL too. IBM even provided a macro; $MVCL. But then, when you look at MVCL usage closely, there are a few defining characteristics that are very useful. It can zero memory, and the millicode (IBM's word for microcode) recognizes 4K boundaries for 4K lengths and optimises it; it's faster than 16 MVCs. There's even a MVPG instruction for moving 4K aligned pages! What are those crazy instruction set designers thinking? The answer's a bit more than just code density; it never really was about that. In all the years I wrote IBM BAL, I never gave code density a serious thought -- with one exception. That was the 4K base address limit; a base register could only span 4K, so code that was bigger than that, you had to have either fancy register footwork or waste registers for multiple bases. It was more about giving assembler programmers choice and variety to get the best out of the box before the advent of optimising compilers; a way, if you like, of exposing the potential of the micro/millicode through the instruction set. "Here I want you to zero memory" meant an MVCL. "Here I am moving 8 bytes from A to B" meant using MVC. A knowledgeable assembler programmer could out-perform a compiler. (Nowadays quality compilers do a much better job of instruction selection than humans, especially for pipelined processors that stall.) Hence CISC instruction sets (at least, IMHO and for IBM). They were there for people and performance, not for code density.
Reply by ●April 4, 20132013-04-04
On 4/4/2013 4:38 AM, Syd Rumpo wrote:> On 29/03/2013 21:00, rickman wrote: >> I have been working with stack based MISC designs in FPGAs for some >> years. All along I have been comparing my work to the work of others. >> These others were the conventional RISC type processors supplied by the >> FPGA vendors as well as the many processor designs done by individuals >> or groups as open source. > > <snip> > > Can you achieve as fast interrupt response times on a register-based > machine as a stack machine? OK, shadow registers buy you one fast > interrupt, but that's sort of a one-level 2D stack. > > Even the venerable RTX2000 had an impressive (IIRC) 200ns interrupt > response time.That's an interesting question. The short answer is yes, but it requires that I provide circuitry to do two things, one is to save both the Processor Status Word (PSW) and the return address to the stack in one cycle. The stack computer has two stacks and I can save these two items in one clock cycle. Currently my register machine uses a stack in memory pointed to by a register, so it would require *two* cycles to save two words. But the memory is dual ported and I can use a tiny bit of extra logic to save both words at once and bump the pointer by two. The other task is to save registers. The stack design doesn't really need to do that, the stack is available for new work and the interrupt routine just needs to finish with the stack in the same state as when it started. I've been thinking about how to handle this in the register machine. The registers are really two registers and one bank of registers. R6 and R7 are "special" in that they have a separate incrementer to support the addressing modes. They need a separate write port so they can be updated in parallel with the other registers. I have considered "saving" the registers by just bumping the start address of the registers in the RAM, but that only saves R0-R5. I could use LUT RAM for R6 and R6 as well. This would provide two sets of registers for R0-R5 and up to 16 sets for R6 and R7. The imbalance isn't very useful, but at least there would be a set for the main program and a set for interrupts with the caveat that nothing can be retained between interrupts. This also means interrupts can't be interrupted other than at specific points where the registers are not used for storage. I'm also going to look at using a block RAM for the registers. With only two read and write ports this makes the multiply step cycle longer though. Once that issue is resolved the interrupt response then becomes the same as the stack machine - 1 clock cycle or 20 ns. -- Rick
Reply by ●April 4, 20132013-04-04
On Apr 3, 5:34=A0pm, "Rod Pemberton" <do_not_h...@notemailnotq.cpm> wrote:> =A0CISC was > typically little-endian to reduce the space needed for integer > encodings.As long as you discount IBM mainframes. They are big endian. Or Borroughs/Unisys; they were big endian too. Or the Motorola 68K; it was big endian. For little-endian CISC, only the VAX and x86 come to mind. Of those only the x86 survives.
Reply by ●April 4, 20132013-04-04
In comp.arch.fpga Alex McDonald <blog@rivadpm.com> wrote: (snip, someone wrote)>> >>> But, of course, this is a fallacy. The same goal is accomplished by >> >>> macro's, and better. Code densitity is the only valid reason.>> Albert, do you have a reference about this?> Let's take two commonly used S/360 opcodes as an example of CISC; some > move operations. MVC (move 0 to 255 bytes) MVCL (move 0 to 16M bytes).MVC moves 1 to 256 bytes, conveniently. (Unless you want 0.)> MVC does no padding or truncation. MVCL can pad and truncate, but > unlike MVC will do nothing and report overflow if the operands > overlap. MVC appears to other processors as a single indivisible > operation; every processor (including IO processors) sees storage as > either before the MVC or after it;I haven't looked recently, but I didn't think it locked out I/O. Seems that one of the favorite tricks for S/360 was modifying channel programs while they are running. (Not to mention self- modifying channel prorams.) Seems that MVC would be convenient for that. It might be that MVC interlocks on CCW fetch such that only whole CCWs are fetched, though.> it's not interruptible. MVCL is > interruptible, and partial products can be observed by other > processors. MVCL requires 4 registers and their contents are updated > after completion of the operation; MVC requires 1 for variable length > moves, 0 for fixed and its contents are preserved. MVCL has a high > code setup cost; MVC has none.> Writing a macro to do multiple MVCs and mimic the behaviour of MVCL? > Why not? It's possible, if a little tricky. And by all accounts, MVC > in a loop is faster than MVCL too. IBM even provided a macro; $MVCL.> But then, when you look at MVCL usage closely, there are a few > defining characteristics that are very useful. It can zero memory, and > the millicode (IBM's word for microcode) recognizes 4K boundaries for > 4K lengths and optimises it; it's faster than 16 MVCs.As far as I understand, millicode isn't exactly like microcode, but does allow for more complicated new instructions to be more easily implemented.> There's even a MVPG instruction for moving 4K aligned pages! What are > those crazy instruction set designers thinking?> The answer's a bit more than just code density; it never really was > about that. In all the years I wrote IBM BAL, I never gave code > density a serious thought -- with one exception. That was the 4K base > address limit; a base register could only span 4K, so code that was > bigger than that, you had to have either fancy register footwork or > waste registers for multiple bases.Compared to VAX, S/360 is somewhat RISCy. Note only three different instruction lengths and, for much of the instruction set only two address modes. If processors fast path the more popular instructions, like L and even MVC, it isn't so far from RISC.> It was more about giving assembler programmers choice and variety to > get the best out of the box before the advent of optimising compilers;Though stories are that even the old Fortran H could come close to good assembly programmers, and likely better than the average assembly programmer.> a way, if you like, of exposing the potential of the micro/millicode > through the instruction set. "Here I want you to zero memory" meant an > MVCL. "Here I am moving 8 bytes from A to B" meant using MVC. A > knowledgeable assembler programmer could out-perform a compiler. > (Nowadays quality compilers do a much better job of instruction > selection than humans, especially for pipelined processors that > stall.)For many processors, MVC was much faster on appropriately aligned data, such as the 8 bytes from A to B. Then again, some might use LD and STD.> Hence CISC instruction sets (at least, IMHO and for IBM). They were > there for people and performance, not for code density.I noticed some time ago that the hex opcodes for add instructions end in A, and for divide in D. (That leaves B for subtract and C for multiply, but not so hard to remember.) If they really wanted to reduce code size, they should have added a load indirect register instruction. (RR format.) A good fraction of L (load) instructions have both base and offset zero, (or, equivalently, index and offset). -- glen





