Java正则表达式非贪婪不起作用

时间:2012-02-14 05:02:48

标签: java regex-greedy

我正在尝试从以下HTML源代码中提取包含字符串“Active”的Item。请注意,活动项目可能会先行。

这是我正在使用的REGEX(试图非贪婪):

<TR class=\\w*?Item>.*?Active.*?</TD></TR>

使用Pattern.CASE_INSENSITIVE | Pattern.DOTALL | Pattern.MULTILINE

它提取整个源代码而不是第二个:...

HTML源代码:

<TR class=Item>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>07D8F15</A> </TD>
<TD style="WIDTH: 150px">IBM 8204-E8A</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - ML 12</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER6</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-D-G13</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel1SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel2SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel3SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2011-04-15</TD>
<TD style="WIDTH: 100px">Cool Down </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_btnDeleteServer disabled>Delete</A> </TD></TR>
<TR class=AlternateItem>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>105ABCD</A> </TD>
<TD style="WIDTH: 150px">IBM Power 770</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - TL 12 SP 01</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER7</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-O-G11</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel1SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel2SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel3SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2012-02-09</TD>
<TD style="WIDTH: 100px">Active </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_btnDeleteServer disabled>Delete</A> </TD></TR>

我们将非常感谢您的帮助!

2 个答案:

答案 0 :(得分:1)

*?zero or more times。您可能需要<TR class=\\w+?Item>

答案 1 :(得分:1)

您可以按Pattern p = Pattern.compile("....", Pattern.MULTILINE);启用多行模式 您可能需要在模式中将</TD></TR>更改为</TD>.*?</TR>