FPGARelated.com
Forums

Please help, Xilinx FIFO problem!

Started by Antti December 21, 2009
On Dec 21, 1:12=A0pm, Antti <antti.luk...@googlemail.com> wrote:
> On Dec 21, 10:21=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > > > On Dec 21, 11:58=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > On Dec 21, 9:50=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > On Dec 21, 9:30=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > > On Dec 21, 7:20=A0pm, Ed McGettigan <ed.mcgetti...@xilinx.com> wr=
ote:
> > > > > > > On Dec 21, 3:01=A0am, Antti <antti.luk...@googlemail.com> wrote=
:
> > > > > > > > On Dec 21, 12:56=A0pm, Symon <symon_bre...@hotmail.com> wrote=
:
> > > > > > > > > Antti wrote: > > > > > > > > > > Xilinx Coregen FIFO, dual clock, most options disable, on=
ly FULL EMPTY
> > > > > > > > > flags present. > > > > > > > > > > signals at input correct, as expected (checked with ChipS=
cope)
> > > > > > > > > signals at output: > > > > > > > > > - double value > > > > > > > > > - missing 1, 2 or 3 values > > > > > > > > > - FIFO will read out random number of OLD entries, this c=
ould be 4
> > > > > > > > > values, or 50% of the FIFO old values > > > > > > > > > I know you will have read this. > > > > > > > > > Can you think of any reason why the Xilinx work-around woul=
dn't work
> > > > > > > > because of your specific implementation? It seems to have d=
ifferent
> > > > > > > > work-arounds depending on whether the read clock is faster =
or slower
> > > > > > > > than the write clock. Do your clocks change frequency? > > > > > > > > > Are you sure your clocks don't have any glitches? The reset=
also?
> > > > > > > > Power's OK? Is your office made of Cobalt 60? > > > > > > > > > HTH., Syms. > > > > > > > > 1) I entered the clock figures in FIFO16 implementationm, but=
the
> > > > > > > error also happens with BRAM based FIFO that do not need work=
arounds
> > > > > > > 2) Clocks DO NOT CHANGE ever, one is MGT recovered clock 125M=
Hz write,
> > > > > > > one is PLB clock 62.5MHz read > > > > > > > 3) Power OK? Well the problem happens at 2 different sites, h=
m yes it
> > > > > > > could be still be power problem > > > > > > > > 4) My office is not of Cobalt 60, ... and its cold here too > > > > > > > > Antti- Hide quoted text - > > > > > > > > - Show quoted text - > > > > > > > Are you sure that this is a FIFO issue and not something else? =
=A0Some
> > > > > > things to think about. > > > > > > > 1) The recovered clock from the MGT is a bit noisy as it moves =
as the
> > > > > > CDR moves. =A0Why are you using this instead of the REFCLK sour=
ce?
> > > > > > > 2) It seems like you have a PLB core that is reading from the F=
IFO,
> > > > > > could the problem be in this? > > > > > > > Ed McGettigan > > > > > > -- > > > > > > Xilinx Inc. > > > > > > Well the MGT datapath and clock system is not done by me, and the=
guy
> > > > > says it is OK all the way it is connected. > > > > > > yes, It is very unlikely to belive that all THREE types of corege=
n
> > > > > FIFO's fail with about same symptoms, but in all > > > > > 3 cased Chipscope sees correct data into fifo, and trash coming o=
ut
> > > > > > the system can span up to 100 boards, all synced to master unit, =
the
> > > > > local refclk is not fully sync to the clock of > > > > > the master unit, so I see no way to use this clock to syncronise =
the
> > > > > fifo? > > > > > > Antti > > > > > PS I just received a attempt to collect the reward, by using non > > > > > xilinx FIFO implementation, i let you all know > > > > > the test results > > > > > Antti > > > > If I remember right (I am no longer at Xilinx) the FIFO is NOT > > > > designed for unequal data width of write and read. (Reason: possibl=
e
> > > > ambiguity of Full and EMPTY) > > > > Since you use two clocks that are roughly 2:1 in frequency, I hope > > > > that you do not try to have double width on one of the ports. > > > > The FIFO must have the same width on both ports. You must design th=
e
> > > > width conversion outside the FIFO. That little circuit will be > > > > synchronous and thus quite simple. > > > > Peter Alfke > > > > well the FIFO is 9b in 9b out so it should work? > > > at least this is what i hoped... > > > > we did not suspect the FIFO as problem at first > > > so spent LOT of time looking for the problem AROUND the FIFOS > > > but.. at least based on what i can see from CS snapshots on fifo > > > inputs and outputs, the only explanation i have is that the FIFO > > > are just goind mad, > > > > of course one option is that its me doing, but i have someone > > > who is in better shape looking over the code as well, and he > > > sees no issues there either. I know the FIFOs should work > > > so there must be some explanation, but so far failing to see it. > > > > Antti > > > PS thank you Peter for the response > > > OK, Antti, > > so you have the same port width, but one clock is about twice as fast > > as the other. > > How do you stop the 125 MHz write clock from filling up the FIFO, > > since you read at only 62 MHz ? > > I hope you are not gating the clock, but rather run it continuously > > and use WE to stop the writing. > > Yes, many of these suggestions are well below your level, but stupid > > problems need stupid investigations. > > Cheers > > Peter > > I am level below ground right now the project is just driving me nuts. > slowly. > To work for months, and end up with Xilinx saying: > The man who could have helped you, left Xilinx last friday. Your > situation is unsupportable. > Well we got out of that situation. > To end up in the new ones. > > The FIFO is never over filled by design. > The fiber link is 99% IDLE sending usually only short 10byte packets > over the link. > > For tesing I generate 10 byte pakets with MOUSE so 1 per second so > there is no doubt > the FIFO is never near full at all. > > Last results: > - ALL 3 types of Xilinx FIFO's same style of errors, about same error > rate > - VHDL FIFO send by CAF reader, uses gray counters, about TEN TIMES > LESS errors then Xilinx implementation, but still all different types > of error did occour: missing values, and FIFO outputtin large junk of > OLD values, that is read pointer changing by some random value > > again, I did not design the MGT clocking and the overall MGT > subsystem, the people who did are either unreachable or unable to > provide any help beyound saying that the implementation (connection of > the FIFO) is done properly. It is also what I have figured out so far, > but.. well somewhere must be problem. > > Antti- Hide quoted text - > > - Show quoted text -
Since you can't get further on the MGT clocking circuit topology, what about the 62.5 MHz read clock? How is this generated? It sounds like it could be glitching. In one of your other posts, you had mention that ChipScope had shown that the write data was correct and that the read data wasn't. Did you have two separate ILA cores with the 125 MHz and 62.5 MHz clocks when you did this testing? Ed McGettigan -- Xilinx Inc
On Dec 21, 4:43=A0pm, Antti <antti.luk...@googlemail.com> wrote:
> On Dec 21, 11:29=A0pm, n...@puntnl.niks (Nico Coesel) wrote: > > > > > Antti <antti.luk...@googlemail.com> wrote: > > >On Dec 21, 3:21=3DA0pm, John McCaskill <jhmccask...@gmail.com> wrote: > > >> On Dec 21, 5:42=3DA0am, Antti <antti.luk...@googlemail.com> wrote: > > > >> > On Dec 21, 1:29=3DA0pm, "maxascent" <maxasc...@yahoo.co.uk> wrote: > > > >> > > >On Dec 21, 12:32=3D3DA0pm, "maxascent" <maxasc...@yahoo.co.uk> =
wrote:
> > >> > > >> Well once you have written and tested your own fifo then you =
would=3D
> > > have > > >> > > i=3D3D > > >> > > >t > > >> > > >> for any other project. It seems like you have wasted a lot of=
time
> > >> > > alread=3D3D > > >> > > >y > > >> > > >> trying to fix the Xilinx version so I dont see that you have =
anyth=3D
> > >ing > > >> > > to > > >> > > >> loose by creating your own. > > > >> > > >> Jon =3D3DA0 =3D3DA0 =3D3DA0 =3D3DA0 > > > >> > > >If you REALLY need todo something else, when your time is at ab=
solut=3D
> > >e > > >> > > >premium > > >> > > >And if the system working (except occasional errors about 2 of =
fiber
> > >> > > >packets are corrupt) > > >> > > >Then you do not go replacing Xilinx validated FIFO solutions wi=
th yo=3D
> > >ur > > >> > > >own, if there are other options. > > > >> > > >If 2 completly different FIFO implementations both have same er=
ror?
> > >> > > >you think 3rd one would instantly work? Could be, yes. > > > >> > > >Antti > > > >> > > In my opinion people tend to use coregen far too often. Looking =
throu=3D
> > >gh > > >> > > some of Xilinx code it is awfull. I went down the route of writi=
ng my=3D
> > > own > > >> > > fifos not because I had a problem with Xilinx fifos but because =
I bel=3D
> > >ieve a > > >> > > fifo written by myself is a lot more flexible and simulates fast=
er th=3D
> > >an the > > >> > > Xilinx version. I also know to as good a degree as I can test th=
at it=3D
> > > will > > >> > > work 100%. > > >> > > I dont really think you can say that their fifos have been valid=
ated =3D
> > >100% > > >> > > if they have to release patches for them. > > > >> > > Jon =3DA0 =3DA0 =3DA0 =3DA0 > > > >> > Dear Jon, > > > >> > I do not feel to be in health right now to write this fifo, so her=
e is
> > >> > the deal: > > > >> > =3DA0 component mgt_fifo > > >> > =3DA0 =3DA0 port ( > > >> > =3DA0 =3DA0 =3DA0 din =3DA0 =3DA0: in =3DA0std_logic_vector(8 down=
to 0);
> > >> > =3DA0 =3DA0 =3DA0 rd_clk : in =3DA0std_logic; > > >> > =3DA0 =3DA0 =3DA0 rd_en =3DA0: in =3DA0std_logic; > > >> > =3DA0 =3DA0 =3DA0 rst =3DA0 =3DA0: in =3DA0std_logic; > > >> > =3DA0 =3DA0 =3DA0 wr_clk : in =3DA0std_logic; > > >> > =3DA0 =3DA0 =3DA0 wr_en =3DA0: in =3DA0std_logic; > > >> > =3DA0 =3DA0 =3DA0 dout =3DA0 : out std_logic_vector(8 downto 0); > > >> > =3DA0 =3DA0 =3DA0 empty =3DA0: out std_logic; > > >> > =3DA0 =3DA0 =3DA0 full =3DA0 : out std_logic); > > >> > =3DA0 end component; > > > >> > if you can write fifo that i can "drop in" and the Xilinx FIFO err=
or
> > >> > is gone, > > >> > then i will stand up, go to postal office and send you 1000 EUR by > > >> > western union. > > >> > If 1000 EUR is not enough, name your price, i will consider it. > > >> > there is no price on the health of our family > > > >> > condition is: DROP IN, WORKS, if i need to troubleshoot, then no p=
ay.
> > > >> > Antti > > > >> Hello Antti, > > > >> If you want to try a different implementation of a FIFO, you can get > > >> the one that the FSL bus uses out of the EDK pcores directory at C: > > >> \Xilinx\11.1\EDK\hw\XilinxProcessorIPLib\pcores\fsl_v20_v2_11_a\hdl > > >> \vhdl. > > > >> There are multiple implementations, including an async BRAM based on=
e
> > >> that has the same ports as you list above, except that it uses exist > > >> instead of empty on the read port. > > > >> That said, I don't expect a third implementation to work instantly > > >> when the previous two implementations had the same error. =3DA0This =
FIFO
> > >> has the full source to it, so it is straight forward to see how it > > >> works, and add ChipScope to observe what is happening around the tim=
e
> > >> of the error. > > > >> If you have not used it before, FPGA editor has the ability to find =
a
> > >> ChipScope ILA core, and change what is connected to it. That can mak=
e
> > >> it much quicker to follow the trail of clues since you avoid having =
to
> > >> go through a full place and route every time you want to look at > > >> something different. > > > >> Is your 62.5 MHz clock a divided version of the 125 MHz clock? You > > >> mention that the 125 MHz is the recovered clock from the MGT, but > > >> there are other options. =3DA0When we did our GigE interface, we use=
d a
> > >> 125 MHz clock from the MGT, but it was not the recovered clock, but > > >> the local MGT PLL. =3DA0This let us use the same 125 MHz clock for a=
ll
> > >> four GigE interfaces and a PMCD to generate a 62.5 MHz clock that is > > >> phase aligned with the 125 MHz clock. > > > >> Regards, > > > >> John McCaskillwww.FasterTechnology.com > > > >Hi > > > >I have tried all 3 variants possible with coregen, > > >all 3 have similar errors > > > >and no, the clocks are not divided version, the 125MHz comes from > > >master over fiber > > >the master could be 100 hops away, the 62.5mhz is derived from local > > >oscillator > > > >so the frequencier are very close but not synchron > > > >Antti > > >who has to give up, at least for a while :( > > >good advice still welcome, if there is any hope or idea how to fix the > > >issue > > >and yes it could be power supply issue at the end of the day also > > > I always write my own fifo's to keep things simple. I keep a write > > pointer, read pointer and number of elements counter in the domain > > with the highest clock frequency. I don't cross the clock domain > > inside the fifo instead I create an interface which does the clock > > domain crossing. I also use an early full signal (say max. elements -X > > depending on the expected latency). This allows for fast FIFO's (no > > cray code counters) with very little logic. > > > The control logic looks like this: > > > if read then read_ptr++; > > if write then write_ptr++; > > if (read=3Dtrue and write=3Dfalse) num_elements--; > > if (write=3Dtrue and read=3Dfalse) num_elements++; > > > if (num_elements>=3D(MAX_ELEMENTS-X)) full=3Dtrue; else full=3Dfalse; > > if (num_elements=3D=3D0) empty=3Dtrue; > > > The external logic should prohibit itself from reading/writing fifo > > when its empty or full. > > > Besides: could your problem be a timing constraint problem? Did you > > specify the amount of time signals may travel from one clock domain to > > the other? The Xilinx tools are not doing this automatically! > > > -- > > Failure does not prove something is impossible, failure simply > > indicates you are not using the right tools... > > =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0"If it doesn't fit, use a bi=
gger hammer!"
> > -------------------------------------------------------------- > > hi > > I was already thinking of writing "simplified FIFO" that is would > work under the conditions it is used, the read is done by PPC software > polling so never too often > > well the clock domains are fully async, so the clock edges of the read- > write > can have any phase they like > > so I assumed if the read and write clock are constrained then it is > enough? > > Antti
Sometimes the simpler things can get in the way of complex issues. Are you certain your read enable and write enables are showing up relative to the correct data? It seems some people expect the read enable to indicate the valid data is being removed from the FIFO while others believe the read enable should produce valid data on the following clock. Double check where the documentation says the valid data should be relative to the enable pulse especially for the read, but check the write as well. ___ How deep do you want your FIFO? Is latency an issue? Do you want rd_en to indicate you're taking valid data or that the next clock is valid? You want wr_en to be present in the same clock cycle as the din, right? Long time no post (partly because I miss having a real newsreader), - John_H
On Dec 22, 1:12=A0am, Ed McGettigan <ed.mcgetti...@xilinx.com> wrote:
> On Dec 21, 1:12=A0pm, Antti <antti.luk...@googlemail.com> wrote: > > > > > > > On Dec 21, 10:21=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > On Dec 21, 11:58=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > On Dec 21, 9:50=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > > On Dec 21, 9:30=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > > > On Dec 21, 7:20=A0pm, Ed McGettigan <ed.mcgetti...@xilinx.com> =
wrote:
> > > > > > > > On Dec 21, 3:01=A0am, Antti <antti.luk...@googlemail.com> wro=
te:
> > > > > > > > > On Dec 21, 12:56=A0pm, Symon <symon_bre...@hotmail.com> wro=
te:
> > > > > > > > > > Antti wrote: > > > > > > > > > > > Xilinx Coregen FIFO, dual clock, most options disable, =
only FULL EMPTY
> > > > > > > > > > flags present. > > > > > > > > > > > signals at input correct, as expected (checked with Chi=
pScope)
> > > > > > > > > > signals at output: > > > > > > > > > > - double value > > > > > > > > > > - missing 1, 2 or 3 values > > > > > > > > > > - FIFO will read out random number of OLD entries, this=
could be 4
> > > > > > > > > > values, or 50% of the FIFO old values > > > > > > > > > > I know you will have read this. > > > > > > > > > > Can you think of any reason why the Xilinx work-around wo=
uldn't work
> > > > > > > > > because of your specific implementation? It seems to have=
different
> > > > > > > > > work-arounds depending on whether the read clock is faste=
r or slower
> > > > > > > > > than the write clock. Do your clocks change frequency? > > > > > > > > > > Are you sure your clocks don't have any glitches? The res=
et also?
> > > > > > > > > Power's OK? Is your office made of Cobalt 60? > > > > > > > > > > HTH., Syms. > > > > > > > > > 1) I entered the clock figures in FIFO16 implementationm, b=
ut the
> > > > > > > > error also happens with BRAM based FIFO that do not need wo=
rkarounds
> > > > > > > > 2) Clocks DO NOT CHANGE ever, one is MGT recovered clock 12=
5MHz write,
> > > > > > > > one is PLB clock 62.5MHz read > > > > > > > > 3) Power OK? Well the problem happens at 2 different sites,=
hm yes it
> > > > > > > > could be still be power problem > > > > > > > > > 4) My office is not of Cobalt 60, ... and its cold here too > > > > > > > > > Antti- Hide quoted text - > > > > > > > > > - Show quoted text - > > > > > > > > Are you sure that this is a FIFO issue and not something else=
? =A0Some
> > > > > > > things to think about. > > > > > > > > 1) The recovered clock from the MGT is a bit noisy as it move=
s as the
> > > > > > > CDR moves. =A0Why are you using this instead of the REFCLK so=
urce?
> > > > > > > > 2) It seems like you have a PLB core that is reading from the=
FIFO,
> > > > > > > could the problem be in this? > > > > > > > > Ed McGettigan > > > > > > > -- > > > > > > > Xilinx Inc. > > > > > > > Well the MGT datapath and clock system is not done by me, and t=
he guy
> > > > > > says it is OK all the way it is connected. > > > > > > > yes, It is very unlikely to belive that all THREE types of core=
gen
> > > > > > FIFO's fail with about same symptoms, but in all > > > > > > 3 cased Chipscope sees correct data into fifo, and trash coming=
out
> > > > > > > the system can span up to 100 boards, all synced to master unit=
, the
> > > > > > local refclk is not fully sync to the clock of > > > > > > the master unit, so I see no way to use this clock to syncronis=
e the
> > > > > > fifo? > > > > > > > Antti > > > > > > PS I just received a attempt to collect the reward, by using no=
n
> > > > > > xilinx FIFO implementation, i let you all know > > > > > > the test results > > > > > > Antti > > > > > If I remember right (I am no longer at Xilinx) the FIFO is NOT > > > > > designed for unequal data width of write and read. (Reason: possi=
ble
> > > > > ambiguity of Full and EMPTY) > > > > > Since you use two clocks that are roughly 2:1 in frequency, I hop=
e
> > > > > that you do not try to have double width on one of the ports. > > > > > The FIFO must have the same width on both ports. You must design =
the
> > > > > width conversion outside the FIFO. That little circuit will be > > > > > synchronous and thus quite simple. > > > > > Peter Alfke > > > > > well the FIFO is 9b in 9b out so it should work? > > > > at least this is what i hoped... > > > > > we did not suspect the FIFO as problem at first > > > > so spent LOT of time looking for the problem AROUND the FIFOS > > > > but.. at least based on what i can see from CS snapshots on fifo > > > > inputs and outputs, the only explanation i have is that the FIFO > > > > are just goind mad, > > > > > of course one option is that its me doing, but i have someone > > > > who is in better shape looking over the code as well, and he > > > > sees no issues there either. I know the FIFOs should work > > > > so there must be some explanation, but so far failing to see it. > > > > > Antti > > > > PS thank you Peter for the response > > > > OK, Antti, > > > so you have the same port width, but one clock is about twice as fast > > > as the other. > > > How do you stop the 125 MHz write clock from filling up the FIFO, > > > since you read at only 62 MHz ? > > > I hope you are not gating the clock, but rather run it continuously > > > and use WE to stop the writing. > > > Yes, many of these suggestions are well below your level, but stupid > > > problems need stupid investigations. > > > Cheers > > > Peter > > > I am level below ground right now the project is just driving me nuts. > > slowly. > > To work for months, and end up with Xilinx saying: > > The man who could have helped you, left Xilinx last friday. Your > > situation is unsupportable. > > Well we got out of that situation. > > To end up in the new ones. > > > The FIFO is never over filled by design. > > The fiber link is 99% IDLE sending usually only short 10byte packets > > over the link. > > > For tesing I generate 10 byte pakets with MOUSE so 1 per second so > > there is no doubt > > the FIFO is never near full at all. > > > Last results: > > - ALL 3 types of Xilinx FIFO's same style of errors, about same error > > rate > > - VHDL FIFO send by CAF reader, uses gray counters, about TEN TIMES > > LESS errors then Xilinx implementation, but still all different types > > of error did occour: missing values, and FIFO outputtin large junk of > > OLD values, that is read pointer changing by some random value > > > again, I did not design the MGT clocking and the overall MGT > > subsystem, the people who did are either unreachable or unable to > > provide any help beyound saying that the implementation (connection of > > the FIFO) is done properly. It is also what I have figured out so far, > > but.. well somewhere must be problem. > > > Antti- Hide quoted text - > > > - Show quoted text - > > Since you can't get further on the MGT clocking circuit topology, what > about the 62.5 MHz read clock? =A0How is this generated? =A0It sounds lik=
e
> it could be glitching. > > In one of your other posts, you had mention that ChipScope had shown > that the write data was correct and that the read data wasn't. =A0Did > you have two separate ILA cores with the 125 MHz and 62.5 MHz clocks > when you did this testing? > > Ed McGettigan > -- > Xilinx Inc
Peter, Ed, et others * yes both clock are running all the time * 125MHz is coming from MGT (recovered clock) there is no gating * the 62.5mhz clock is PLB clock directly there is no gating * the 62.5mhz read is generated as edge detect that generates 1 clk wide pulse on PLB reads * I used separate ILA cores in different clock domains * Routing out the 125MHz for external scope would not show the internal signal same as it is seen by the FIFO module, besides the IOB characteristics would filter out something and introduce an delay so the measurement would not be likely to show anything. OTOH the chipscope inside the FPGA also doesnt tell much about the clock, except that the data if clocked with the selected clock is latched properly at the same conditions where FIFO does go crazy. * The design occupies about 80% of all available resources of Virtex-4 FX40, in order to see the error, I have to start 2 units with GbE to the first one and fiber link between the two, and send specific commands to the master unit where the packets are processed by PPC running custom firmware, that then triggers the condition in the slave where then the problem can be seen. If anyone says this kind of system can be simulated with meaningful results I am all ears to know the setup for this. It doesnt make sense to simulate Xilinx FIFO's they are almost certain not to exhibit the observed fault behavior. the possibilities i still see are: 1) one of the clocks has something really "bad" in it, do not even know what it could be: 1.2Ghz ringing? short bursts of some very high frequency that do not trigger CS but do trigger FIFO ? 2) Xilinx tools are missing the timing that badly that all 4 type of fifos inhibit similar error, but parallel connected CS core doesnt? 3) Problem with power supply? 4) Unspecified technical problem? 5) Me needing in sign off from this project to preserve my sanity? It could be 5, it can be that the problem is there but I constantly connect CS to some other clock and the FIFO's are one some other clock that has problem. Well I have asked help, and a fellow engineer has looked over the clock routing and what he has said is that it is all OK the way it is right now. Maybe he needs a break as well. Antti
On Dec 22, 6:40=A0am, John_H <newsgr...@johnhandwork.com> wrote:
> On Dec 21, 4:43=A0pm, Antti <antti.luk...@googlemail.com> wrote: > > > > > > > On Dec 21, 11:29=A0pm, n...@puntnl.niks (Nico Coesel) wrote: > > > > Antti <antti.luk...@googlemail.com> wrote: > > > >On Dec 21, 3:21=3DA0pm, John McCaskill <jhmccask...@gmail.com> wrote=
:
> > > >> On Dec 21, 5:42=3DA0am, Antti <antti.luk...@googlemail.com> wrote: > > > > >> > On Dec 21, 1:29=3DA0pm, "maxascent" <maxasc...@yahoo.co.uk> wrot=
e:
> > > > >> > > >On Dec 21, 12:32=3D3DA0pm, "maxascent" <maxasc...@yahoo.co.uk= > wrote: > > > >> > > >> Well once you have written and tested your own fifo then yo=
u would=3D
> > > > have > > > >> > > i=3D3D > > > >> > > >t > > > >> > > >> for any other project. It seems like you have wasted a lot =
of time
> > > >> > > alread=3D3D > > > >> > > >y > > > >> > > >> trying to fix the Xilinx version so I dont see that you hav=
e anyth=3D
> > > >ing > > > >> > > to > > > >> > > >> loose by creating your own. > > > > >> > > >> Jon =3D3DA0 =3D3DA0 =3D3DA0 =3D3DA0 > > > > >> > > >If you REALLY need todo something else, when your time is at =
absolut=3D
> > > >e > > > >> > > >premium > > > >> > > >And if the system working (except occasional errors about 2 o=
f fiber
> > > >> > > >packets are corrupt) > > > >> > > >Then you do not go replacing Xilinx validated FIFO solutions =
with yo=3D
> > > >ur > > > >> > > >own, if there are other options. > > > > >> > > >If 2 completly different FIFO implementations both have same =
error?
> > > >> > > >you think 3rd one would instantly work? Could be, yes. > > > > >> > > >Antti > > > > >> > > In my opinion people tend to use coregen far too often. Lookin=
g throu=3D
> > > >gh > > > >> > > some of Xilinx code it is awfull. I went down the route of wri=
ting my=3D
> > > > own > > > >> > > fifos not because I had a problem with Xilinx fifos but becaus=
e I bel=3D
> > > >ieve a > > > >> > > fifo written by myself is a lot more flexible and simulates fa=
ster th=3D
> > > >an the > > > >> > > Xilinx version. I also know to as good a degree as I can test =
that it=3D
> > > > will > > > >> > > work 100%. > > > >> > > I dont really think you can say that their fifos have been val=
idated =3D
> > > >100% > > > >> > > if they have to release patches for them. > > > > >> > > Jon =3DA0 =3DA0 =3DA0 =3DA0 > > > > >> > Dear Jon, > > > > >> > I do not feel to be in health right now to write this fifo, so h=
ere is
> > > >> > the deal: > > > > >> > =3DA0 component mgt_fifo > > > >> > =3DA0 =3DA0 port ( > > > >> > =3DA0 =3DA0 =3DA0 din =3DA0 =3DA0: in =3DA0std_logic_vector(8 do=
wnto 0);
> > > >> > =3DA0 =3DA0 =3DA0 rd_clk : in =3DA0std_logic; > > > >> > =3DA0 =3DA0 =3DA0 rd_en =3DA0: in =3DA0std_logic; > > > >> > =3DA0 =3DA0 =3DA0 rst =3DA0 =3DA0: in =3DA0std_logic; > > > >> > =3DA0 =3DA0 =3DA0 wr_clk : in =3DA0std_logic; > > > >> > =3DA0 =3DA0 =3DA0 wr_en =3DA0: in =3DA0std_logic; > > > >> > =3DA0 =3DA0 =3DA0 dout =3DA0 : out std_logic_vector(8 downto 0); > > > >> > =3DA0 =3DA0 =3DA0 empty =3DA0: out std_logic; > > > >> > =3DA0 =3DA0 =3DA0 full =3DA0 : out std_logic); > > > >> > =3DA0 end component; > > > > >> > if you can write fifo that i can "drop in" and the Xilinx FIFO e=
rror
> > > >> > is gone, > > > >> > then i will stand up, go to postal office and send you 1000 EUR =
by
> > > >> > western union. > > > >> > If 1000 EUR is not enough, name your price, i will consider it. > > > >> > there is no price on the health of our family > > > > >> > condition is: DROP IN, WORKS, if i need to troubleshoot, then no=
pay.
> > > > >> > Antti > > > > >> Hello Antti, > > > > >> If you want to try a different implementation of a FIFO, you can g=
et
> > > >> the one that the FSL bus uses out of the EDK pcores directory at C=
:
> > > >> \Xilinx\11.1\EDK\hw\XilinxProcessorIPLib\pcores\fsl_v20_v2_11_a\hd=
l
> > > >> \vhdl. > > > > >> There are multiple implementations, including an async BRAM based =
one
> > > >> that has the same ports as you list above, except that it uses exi=
st
> > > >> instead of empty on the read port. > > > > >> That said, I don't expect a third implementation to work instantly > > > >> when the previous two implementations had the same error. =3DA0Thi=
s FIFO
> > > >> has the full source to it, so it is straight forward to see how it > > > >> works, and add ChipScope to observe what is happening around the t=
ime
> > > >> of the error. > > > > >> If you have not used it before, FPGA editor has the ability to fin=
d a
> > > >> ChipScope ILA core, and change what is connected to it. That can m=
ake
> > > >> it much quicker to follow the trail of clues since you avoid havin=
g to
> > > >> go through a full place and route every time you want to look at > > > >> something different. > > > > >> Is your 62.5 MHz clock a divided version of the 125 MHz clock? You > > > >> mention that the 125 MHz is the recovered clock from the MGT, but > > > >> there are other options. =3DA0When we did our GigE interface, we u=
sed a
> > > >> 125 MHz clock from the MGT, but it was not the recovered clock, bu=
t
> > > >> the local MGT PLL. =3DA0This let us use the same 125 MHz clock for=
all
> > > >> four GigE interfaces and a PMCD to generate a 62.5 MHz clock that =
is
> > > >> phase aligned with the 125 MHz clock. > > > > >> Regards, > > > > >> John McCaskillwww.FasterTechnology.com > > > > >Hi > > > > >I have tried all 3 variants possible with coregen, > > > >all 3 have similar errors > > > > >and no, the clocks are not divided version, the 125MHz comes from > > > >master over fiber > > > >the master could be 100 hops away, the 62.5mhz is derived from local > > > >oscillator > > > > >so the frequencier are very close but not synchron > > > > >Antti > > > >who has to give up, at least for a while :( > > > >good advice still welcome, if there is any hope or idea how to fix t=
he
> > > >issue > > > >and yes it could be power supply issue at the end of the day also > > > > I always write my own fifo's to keep things simple. I keep a write > > > pointer, read pointer and number of elements counter in the domain > > > with the highest clock frequency. I don't cross the clock domain > > > inside the fifo instead I create an interface which does the clock > > > domain crossing. I also use an early full signal (say max. elements -=
X
> > > depending on the expected latency). This allows for fast FIFO's (no > > > cray code counters) with very little logic. > > > > The control logic looks like this: > > > > if read then read_ptr++; > > > if write then write_ptr++; > > > if (read=3Dtrue and write=3Dfalse) num_elements--; > > > if (write=3Dtrue and read=3Dfalse) num_elements++; > > > > if (num_elements>=3D(MAX_ELEMENTS-X)) full=3Dtrue; else full=3Dfalse; > > > if (num_elements=3D=3D0) empty=3Dtrue; > > > > The external logic should prohibit itself from reading/writing fifo > > > when its empty or full. > > > > Besides: could your problem be a timing constraint problem? Did you > > > specify the amount of time signals may travel from one clock domain t=
o
> > > the other? The Xilinx tools are not doing this automatically! > > > > -- > > > Failure does not prove something is impossible, failure simply > > > indicates you are not using the right tools... > > > =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0"If it doesn't fit, use a =
bigger hammer!"
> > > -------------------------------------------------------------- > > > hi > > > I was already thinking of writing "simplified FIFO" that is would > > work under the conditions it is used, the read is done by PPC software > > polling so never too often > > > well the clock domains are fully async, so the clock edges of the read- > > write > > can have any phase they like > > > so I assumed if the read and write clock are constrained then it is > > enough? > > > Antti > > Sometimes the simpler things can get in the way of complex issues. > > Are you certain your read enable and write enables are showing up > relative to the correct data? > It seems some people expect the read enable to indicate the valid data > is being removed from the FIFO while others believe the read enable > should produce valid data on the following clock. > > Double check where the documentation says the valid data should be > relative to the enable pulse especially for the read, but check the > write as well. > ___ > > How deep do you want your FIFO? > Is latency an issue? > Do you want rd_en to indicate you're taking valid data or that the > next clock is valid? > You want wr_en to be present in the same clock cycle as the din, > right? > > Long time no post (partly because I miss having a real newsreader), > - John_H
Hi John, 1 the FIFO is supposed to be SIMPLEST possible MGT receiver, FIFO wr_en is active when the incoming char is not IDLE. 2 Latency is absolutly NO issue, PPC is pulling the data extremly slow anyway :( 3 rd_en almost do not care, well currently it is wrong, 1 clock too late so PPC doesnt pull the last value from fifo (it is pulled when new data comes in), but this minor issue does really not explain the error where the fifo reads out out half of the old values Antti
On Dec 21, 3:12=A0pm, Antti <antti.luk...@googlemail.com> wrote:
> On Dec 21, 10:21=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > On Dec 21, 11:58=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > On Dec 21, 9:50=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > On Dec 21, 9:30=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > > On Dec 21, 7:20=A0pm, Ed McGettigan <ed.mcgetti...@xilinx.com> wr=
ote:
> > > > > > > On Dec 21, 3:01=A0am, Antti <antti.luk...@googlemail.com> wrote=
:
> > > > > > > > On Dec 21, 12:56=A0pm, Symon <symon_bre...@hotmail.com> wrote=
:
> > > > > > > > > Antti wrote: > > > > > > > > > > Xilinx Coregen FIFO, dual clock, most options disable, on=
ly FULL EMPTY
> > > > > > > > > flags present. > > > > > > > > > > signals at input correct, as expected (checked with ChipS=
cope)
> > > > > > > > > signals at output: > > > > > > > > > - double value > > > > > > > > > - missing 1, 2 or 3 values > > > > > > > > > - FIFO will read out random number of OLD entries, this c=
ould be 4
> > > > > > > > > values, or 50% of the FIFO old values > > > > > > > > > I know you will have read this. > > > > > > > > > Can you think of any reason why the Xilinx work-around woul=
dn't work
> > > > > > > > because of your specific implementation? It seems to have d=
ifferent
> > > > > > > > work-arounds depending on whether the read clock is faster =
or slower
> > > > > > > > than the write clock. Do your clocks change frequency? > > > > > > > > > Are you sure your clocks don't have any glitches? The reset=
also?
> > > > > > > > Power's OK? Is your office made of Cobalt 60? > > > > > > > > > HTH., Syms. > > > > > > > > 1) I entered the clock figures in FIFO16 implementationm, but=
the
> > > > > > > error also happens with BRAM based FIFO that do not need work=
arounds
> > > > > > > 2) Clocks DO NOT CHANGE ever, one is MGT recovered clock 125M=
Hz write,
> > > > > > > one is PLB clock 62.5MHz read > > > > > > > 3) Power OK? Well the problem happens at 2 different sites, h=
m yes it
> > > > > > > could be still be power problem > > > > > > > > 4) My office is not of Cobalt 60, ... and its cold here too > > > > > > > > Antti- Hide quoted text - > > > > > > > > - Show quoted text - > > > > > > > Are you sure that this is a FIFO issue and not something else? =
=A0Some
> > > > > > things to think about. > > > > > > > 1) The recovered clock from the MGT is a bit noisy as it moves =
as the
> > > > > > CDR moves. =A0Why are you using this instead of the REFCLK sour=
ce?
> > > > > > > 2) It seems like you have a PLB core that is reading from the F=
IFO,
> > > > > > could the problem be in this? > > > > > > > Ed McGettigan > > > > > > -- > > > > > > Xilinx Inc. > > > > > > Well the MGT datapath and clock system is not done by me, and the=
guy
> > > > > says it is OK all the way it is connected. > > > > > > yes, It is very unlikely to belive that all THREE types of corege=
n
> > > > > FIFO's fail with about same symptoms, but in all > > > > > 3 cased Chipscope sees correct data into fifo, and trash coming o=
ut
> > > > > > the system can span up to 100 boards, all synced to master unit, =
the
> > > > > local refclk is not fully sync to the clock of > > > > > the master unit, so I see no way to use this clock to syncronise =
the
> > > > > fifo? > > > > > > Antti > > > > > PS I just received a attempt to collect the reward, by using non > > > > > xilinx FIFO implementation, i let you all know > > > > > the test results > > > > > Antti > > > > If I remember right (I am no longer at Xilinx) the FIFO is NOT > > > > designed for unequal data width of write and read. (Reason: possibl=
e
> > > > ambiguity of Full and EMPTY) > > > > Since you use two clocks that are roughly 2:1 in frequency, I hope > > > > that you do not try to have double width on one of the ports. > > > > The FIFO must have the same width on both ports. You must design th=
e
> > > > width conversion outside the FIFO. That little circuit will be > > > > synchronous and thus quite simple. > > > > Peter Alfke > > > > well the FIFO is 9b in 9b out so it should work? > > > at least this is what i hoped... > > > > we did not suspect the FIFO as problem at first > > > so spent LOT of time looking for the problem AROUND the FIFOS > > > but.. at least based on what i can see from CS snapshots on fifo > > > inputs and outputs, the only explanation i have is that the FIFO > > > are just goind mad, > > > > of course one option is that its me doing, but i have someone > > > who is in better shape looking over the code as well, and he > > > sees no issues there either. I know the FIFOs should work > > > so there must be some explanation, but so far failing to see it. > > > > Antti > > > PS thank you Peter for the response > > > OK, Antti, > > so you have the same port width, but one clock is about twice as fast > > as the other. > > How do you stop the 125 MHz write clock from filling up the FIFO, > > since you read at only 62 MHz ? > > I hope you are not gating the clock, but rather run it continuously > > and use WE to stop the writing. > > Yes, many of these suggestions are well below your level, but stupid > > problems need stupid investigations. > > Cheers > > Peter > > I am level below ground right now the project is just driving me nuts. > slowly. > To work for months, and end up with Xilinx saying: > The man who could have helped you, left Xilinx last friday. Your > situation is unsupportable. > Well we got out of that situation. > To end up in the new ones. > > The FIFO is never over filled by design. > The fiber link is 99% IDLE sending usually only short 10byte packets > over the link. > > For tesing I generate 10 byte pakets with MOUSE so 1 per second so > there is no doubt > the FIFO is never near full at all. > > Last results: > - ALL 3 types of Xilinx FIFO's same style of errors, about same error > rate > - VHDL FIFO send by CAF reader, uses gray counters, about TEN TIMES > LESS errors then Xilinx implementation, but still all different types > of error did occour: missing values, and FIFO outputtin large junk of > OLD values, that is read pointer changing by some random value > > again, I did not design the MGT clocking and the overall MGT > subsystem, the people who did are either unreachable or unable to > provide any help beyound saying that the implementation (connection of > the FIFO) is done properly. It is also what I have figured out so far, > but.. well somewhere must be problem. > > Antti
Hello Antti, With four different FIFOs all failing, it is not likely that they are the source of the problem, just where the symptoms are showing up, as if you did not already know that. If you still want suggestions, here are a few. First, I always consider having an error condition I can trigger on to be worth its weight in gold and you apparently have one in the FIFO. Put in ChipScope with multiple ILAs observing one of the FIFOs that you have source code for. Use what ever you are currently triggering on to trigger the other ILAs. Put one on the write clock domain, and one on the read clock domain. Have them look at all of the IOs, as well as the counters and other logic in the FIFO. I doubt that you will find a problem with the FIFO, but something will look wrong and give you a clue to follow. Also use separate ILAs to watch the read and write clocks. I am always suspicious of IO clocks, I have seen too many problems with them. If one of those clocks is having a problem, and you are using that clock as the clock for the ILA, you will not see the clock problem with that ILA. Since you are using the recovered clock instead of the reference clock (which you can do, and is how we do it), I would pay extra attention to it. Over sample the read and write clocks by either using one faster clock, or multiple ILAs running on multiple phases of a faster clock. On a Virtex-4FX, we have multiple MGTs/EMACs running GigE. We use the 125 MHz reference clock instead of the recovered clock so we only have one 125 MHz clock to deal with. We feed it through a PMCD to generate the 62.5 MHz clock so that they are not asynchronous. That give us a bit less to have to deal with. Do you have access to a digital storage oscilloscope? If so, run the ILA trigger out of the FPGA and use that to trigger the scope. Use it to look at the clocks and power supplies, and anything else that the other test turned up. Use the timing analyzer to look for unconstrained paths. Look for any cross clock domain buses that have more than a cycle of skew on them. I have not seen that cause problems yet, but I use from to constraints to minimize skew to prevent a gray coded bus from having more than a cycle of skew crossing domains and causing problems. I don't think it is a high probability, but your symptoms remind me of the time we wrote our own FIFO that had different read and write widths and incremented the Gray code counter by two. That would cause two bits to change at a time, and eventually that would cause it to fail. Good luck, and remember that it it was easy, it would not be called hardware. John McCaskill www.FasterTechnology.com
On Dec 22, 7:02=A0am, John McCaskill <jhmccask...@gmail.com> wrote:
> On Dec 21, 3:12=A0pm, Antti <antti.luk...@googlemail.com> wrote: > > > > > > > On Dec 21, 10:21=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > On Dec 21, 11:58=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > On Dec 21, 9:50=A0pm, Peter Alfke <al...@sbcglobal.net> wrote: > > > > > > On Dec 21, 9:30=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > > > On Dec 21, 7:20=A0pm, Ed McGettigan <ed.mcgetti...@xilinx.com> =
wrote:
> > > > > > > > On Dec 21, 3:01=A0am, Antti <antti.luk...@googlemail.com> wro=
te:
> > > > > > > > > On Dec 21, 12:56=A0pm, Symon <symon_bre...@hotmail.com> wro=
te:
> > > > > > > > > > Antti wrote: > > > > > > > > > > > Xilinx Coregen FIFO, dual clock, most options disable, =
only FULL EMPTY
> > > > > > > > > > flags present. > > > > > > > > > > > signals at input correct, as expected (checked with Chi=
pScope)
> > > > > > > > > > signals at output: > > > > > > > > > > - double value > > > > > > > > > > - missing 1, 2 or 3 values > > > > > > > > > > - FIFO will read out random number of OLD entries, this=
could be 4
> > > > > > > > > > values, or 50% of the FIFO old values > > > > > > > > > > I know you will have read this. > > > > > > > > > > Can you think of any reason why the Xilinx work-around wo=
uldn't work
> > > > > > > > > because of your specific implementation? It seems to have=
different
> > > > > > > > > work-arounds depending on whether the read clock is faste=
r or slower
> > > > > > > > > than the write clock. Do your clocks change frequency? > > > > > > > > > > Are you sure your clocks don't have any glitches? The res=
et also?
> > > > > > > > > Power's OK? Is your office made of Cobalt 60? > > > > > > > > > > HTH., Syms. > > > > > > > > > 1) I entered the clock figures in FIFO16 implementationm, b=
ut the
> > > > > > > > error also happens with BRAM based FIFO that do not need wo=
rkarounds
> > > > > > > > 2) Clocks DO NOT CHANGE ever, one is MGT recovered clock 12=
5MHz write,
> > > > > > > > one is PLB clock 62.5MHz read > > > > > > > > 3) Power OK? Well the problem happens at 2 different sites,=
hm yes it
> > > > > > > > could be still be power problem > > > > > > > > > 4) My office is not of Cobalt 60, ... and its cold here too > > > > > > > > > Antti- Hide quoted text - > > > > > > > > > - Show quoted text - > > > > > > > > Are you sure that this is a FIFO issue and not something else=
? =A0Some
> > > > > > > things to think about. > > > > > > > > 1) The recovered clock from the MGT is a bit noisy as it move=
s as the
> > > > > > > CDR moves. =A0Why are you using this instead of the REFCLK so=
urce?
> > > > > > > > 2) It seems like you have a PLB core that is reading from the=
FIFO,
> > > > > > > could the problem be in this? > > > > > > > > Ed McGettigan > > > > > > > -- > > > > > > > Xilinx Inc. > > > > > > > Well the MGT datapath and clock system is not done by me, and t=
he guy
> > > > > > says it is OK all the way it is connected. > > > > > > > yes, It is very unlikely to belive that all THREE types of core=
gen
> > > > > > FIFO's fail with about same symptoms, but in all > > > > > > 3 cased Chipscope sees correct data into fifo, and trash coming=
out
> > > > > > > the system can span up to 100 boards, all synced to master unit=
, the
> > > > > > local refclk is not fully sync to the clock of > > > > > > the master unit, so I see no way to use this clock to syncronis=
e the
> > > > > > fifo? > > > > > > > Antti > > > > > > PS I just received a attempt to collect the reward, by using no=
n
> > > > > > xilinx FIFO implementation, i let you all know > > > > > > the test results > > > > > > Antti > > > > > If I remember right (I am no longer at Xilinx) the FIFO is NOT > > > > > designed for unequal data width of write and read. (Reason: possi=
ble
> > > > > ambiguity of Full and EMPTY) > > > > > Since you use two clocks that are roughly 2:1 in frequency, I hop=
e
> > > > > that you do not try to have double width on one of the ports. > > > > > The FIFO must have the same width on both ports. You must design =
the
> > > > > width conversion outside the FIFO. That little circuit will be > > > > > synchronous and thus quite simple. > > > > > Peter Alfke > > > > > well the FIFO is 9b in 9b out so it should work? > > > > at least this is what i hoped... > > > > > we did not suspect the FIFO as problem at first > > > > so spent LOT of time looking for the problem AROUND the FIFOS > > > > but.. at least based on what i can see from CS snapshots on fifo > > > > inputs and outputs, the only explanation i have is that the FIFO > > > > are just goind mad, > > > > > of course one option is that its me doing, but i have someone > > > > who is in better shape looking over the code as well, and he > > > > sees no issues there either. I know the FIFOs should work > > > > so there must be some explanation, but so far failing to see it. > > > > > Antti > > > > PS thank you Peter for the response > > > > OK, Antti, > > > so you have the same port width, but one clock is about twice as fast > > > as the other. > > > How do you stop the 125 MHz write clock from filling up the FIFO, > > > since you read at only 62 MHz ? > > > I hope you are not gating the clock, but rather run it continuously > > > and use WE to stop the writing. > > > Yes, many of these suggestions are well below your level, but stupid > > > problems need stupid investigations. > > > Cheers > > > Peter > > > I am level below ground right now the project is just driving me nuts. > > slowly. > > To work for months, and end up with Xilinx saying: > > The man who could have helped you, left Xilinx last friday. Your > > situation is unsupportable. > > Well we got out of that situation. > > To end up in the new ones. > > > The FIFO is never over filled by design. > > The fiber link is 99% IDLE sending usually only short 10byte packets > > over the link. > > > For tesing I generate 10 byte pakets with MOUSE so 1 per second so > > there is no doubt > > the FIFO is never near full at all. > > > Last results: > > - ALL 3 types of Xilinx FIFO's same style of errors, about same error > > rate > > - VHDL FIFO send by CAF reader, uses gray counters, about TEN TIMES > > LESS errors then Xilinx implementation, but still all different types > > of error did occour: missing values, and FIFO outputtin large junk of > > OLD values, that is read pointer changing by some random value > > > again, I did not design the MGT clocking and the overall MGT > > subsystem, the people who did are either unreachable or unable to > > provide any help beyound saying that the implementation (connection of > > the FIFO) is done properly. It is also what I have figured out so far, > > but.. well somewhere must be problem. > > > Antti > > Hello Antti, > > With four different FIFOs all failing, it is not likely that they are > the source of the problem, just where the symptoms are showing up, as > if you did not already know that. > > If you still want suggestions, here are a few. > > First, I always consider having an error condition I can trigger on to > be worth its weight in gold and you apparently have one in the FIFO. > Put in ChipScope with multiple ILAs observing one of the FIFOs that > you have source code for. =A0Use what ever you are currently triggering > on to trigger the other ILAs. =A0Put one on the write clock domain, and > one on the read clock domain. =A0Have them look at all of the IOs, as > well as the counters and other logic in the FIFO. =A0I doubt that you > will find a problem with the FIFO, but something will look wrong and > give you a clue to follow. > > Also use separate ILAs to watch the read and write clocks. =A0I am > always suspicious of IO clocks, I have seen too many problems with > them. If one of those clocks is having a problem, and you are using > that clock as the clock for the ILA, you will not see the clock > problem with that ILA. Since you are using the recovered clock instead > of the reference clock (which you can do, and is how we do it), I > would pay extra attention to it. =A0Over sample the read and write > clocks by either using one faster clock, or multiple ILAs running on > multiple phases of a faster clock. =A0On a Virtex-4FX, we have multiple > MGTs/EMACs running GigE. =A0We use the 125 MHz reference clock instead > of the recovered clock so we only have one 125 MHz clock to deal > with. =A0We feed it through a PMCD to generate the 62.5 MHz clock so > that they are not asynchronous. =A0 That give us a bit less to have to > deal with. > > Do you have access to a digital storage oscilloscope? =A0If so, =A0run th=
e
> ILA trigger out of the FPGA and use that to trigger the scope. Use it > to look at the clocks and power supplies, and anything else that the > other test turned up. > > Use the timing analyzer to look for unconstrained paths. Look for any > cross clock domain buses that have more than a cycle of skew on them. > I have not seen that cause problems yet, but I use from to constraints > to minimize skew to prevent a gray coded bus from having more than a > cycle of skew crossing domains and causing problems. =A0I don't think it > is a high probability, but your symptoms remind me of the time we > wrote our own FIFO that had different read and write widths and > incremented the Gray code counter by two. That would cause two bits to > change at a time, and eventually that would cause it to fail. > > Good luck, and remember that it it was easy, it would not be called > hardware. > > John McCaskillwww.FasterTechnology.com
Thank you John, Antti
Antti <antti.lukats@googlemail.com> wrote:
(big snip)
 
> 1 the FIFO is supposed to be SIMPLEST possible MGT receiver, FIFO > wr_en is active when the incoming char is not IDLE.
> 2 Latency is absolutly NO issue, PPC is pulling the data extremly slow > anyway :(
> 3 rd_en almost do not care, well currently it is wrong, 1 clock too > late so PPC doesnt pull the last value from fifo (it is pulled when > new data comes in), but this minor issue does really not explain the > error where the fifo reads out out half of the old values
If the fifo is empty half the time, then half the time you will be reading the wrong value. -- glen
To follow on, here are some of my thoughts:


- I would try to limit the scope of this issue by using a chain of three async identical FIFOs (with the control signals properly forwarded: the point is to make the whole thing transparent, although with increased cycle latency)
  [MGT] ---> [FIFO #1] ---> [FIFO #2] ---> [FIFO #3] --->  [PLB]
    MGT, FIFO #1 (both ports), inbound port of FIFO #2 @ 125 MHz
    Outbound port of FIFO #2, FIFO #3 (both ports), and PLB @ 62.5 MHz
The FIFO #1 and #3 are useless but they may experience the issue you are facing, bringing up interesting facts such as knowing which port is going south.


Also I guess that you already went through the obvious items:

- I would check and re-check __myself__ the clocking scheme inside and outside the FPGA
- Same thing with power-supply
- Check that all I/O pads are LOC'ed (I once had an unconstrained pad due to a typo inside the UCF file, nasty things followed)
- Check that the FIFO reset is performed correctly (all clock stable, FIFO state is idle) and meets the required duration
- A good sleep, cold shower and breakfast are very effective when dealing with though issues !!


-- 
Matthieu Michon <prenom.nom@gmail.com>
On Dec 22, 12:00=A0am, Antti <antti.luk...@googlemail.com> wrote:
> On Dec 22, 6:40=A0am, John_H <newsgr...@johnhandwork.com> wrote: > > > Sometimes the simpler things can get in the way of complex issues. > > > Are you certain your read enable and write enables are showing up > > relative to the correct data? > > It seems some people expect the read enable to indicate the valid data > > is being removed from the FIFO while others believe the read enable > > should produce valid data on the following clock. > > > Double check where the documentation says the valid data should be > > relative to the enable pulse especially for the read, but check the > > write as well. > > ___ > > > How deep do you want your FIFO? > > Is latency an issue? > > Do you want rd_en to indicate you're taking valid data or that the > > next clock is valid? > > You want wr_en to be present in the same clock cycle as the din, > > right? > > > Long time no post (partly because I miss having a real newsreader), > > - John_H > > Hi John, > > 1 the FIFO is supposed to be SIMPLEST possible MGT receiver, FIFO > wr_en is active when the incoming char is not IDLE. > 2 Latency is absolutly NO issue, PPC is pulling the data extremly slow > anyway :( > 3 rd_en almost do not care, well currently it is wrong, 1 clock too > late so PPC doesnt pull the last value from fifo (it is pulled when > new data comes in), but this minor issue does really not explain the > error where the fifo reads out out half of the old values > > Antti
If the rd_en is one cycle off, the data during that cycle is undefined. If the rd_en is active for two cycles, the data extracted will be precisely one cycle off for the rd_en pulses after the first. The specific FIFO implementation may provide what looks like valid data - or not - during the first of those consecutive rd_en pulses. I would *love* to know how much data is "good" versus "bad" with the rd_en realigned. - John_H
On Dec 22, 8:06=A0am, John_H <newsgr...@johnhandwork.com> wrote:
> On Dec 22, 12:00=A0am, Antti <antti.luk...@googlemail.com> wrote: > > > > > On Dec 22, 6:40=A0am, John_H <newsgr...@johnhandwork.com> wrote: > > > > Sometimes the simpler things can get in the way of complex issues. > > > > Are you certain your read enable and write enables are showing up > > > relative to the correct data? > > > It seems some people expect the read enable to indicate the valid dat=
a
> > > is being removed from the FIFO while others believe the read enable > > > should produce valid data on the following clock. > > > > Double check where the documentation says the valid data should be > > > relative to the enable pulse especially for the read, but check the > > > write as well. > > > ___ > > > > How deep do you want your FIFO? > > > Is latency an issue? > > > Do you want rd_en to indicate you're taking valid data or that the > > > next clock is valid? > > > You want wr_en to be present in the same clock cycle as the din, > > > right? > > > > Long time no post (partly because I miss having a real newsreader), > > > - John_H > > > Hi John, > > > 1 the FIFO is supposed to be SIMPLEST possible MGT receiver, FIFO > > wr_en is active when the incoming char is not IDLE. > > 2 Latency is absolutly NO issue, PPC is pulling the data extremly slow > > anyway :( > > 3 rd_en almost do not care, well currently it is wrong, 1 clock too > > late so PPC doesnt pull the last value from fifo (it is pulled when > > new data comes in), but this minor issue does really not explain the > > error where the fifo reads out out half of the old values > > > Antti > > If the rd_en is one cycle off, the data during that cycle is > undefined. > > If the rd_en is active for two cycles, the data extracted will be > precisely one cycle off for the rd_en pulses after the first. > > The specific FIFO implementation may provide what looks like valid > data - or not - during the first of those consecutive rd_en pulses. > > I would *love* to know how much data is "good" versus "bad" with the > rd_en realigned. > > - John_H
Rereading some stuff, perhaps I wasn't clear: the data for a read pulse may be expected by the system the same clock as the read enable - indicating the design is "taking valid data" from the FIFO on that clock where rd_en and dout are both valid the same clock cycle - while the FIFO expects the dout to be valid the clock *after* the rd_en. Or vice-versa. It's this "which clock cycle" issue between rd_en and the dout corresponding to that rd_en that trips up some engineers.